Ten workers, one provider account
Context
We draft patent sections with a hosted model. The drafting work runs on ten worker processes
behind a load balancer, and all ten use the same provider account, which is capped at 60 requests
and 20,000 tokens a minute.
Since we scaled from three workers to ten, drafting jobs have started failing. The provider is
sending back 429s in bursts, attorneys see sections that never finish, and the jobs that do get
through are not the ones anyone prioritised. Nobody has changed the model code.
Each worker builds its own rate limiter when it starts, and that limiter is sized for the whole
account. Ten workers each politely holding themselves to 60 requests a minute is a fleet asking
for 600.
The provider and the shared store in this repo are deterministic offline stand-ins. There is no
network and no API key.
Your task
Three phases, in order. Phase 3 is a stretch goal. You are not expected to finish all three,
and most people do not reach phase 3 - we care much more about how you work through phases 1
and 2 than about getting to the end.
1. Fix a bug in the code as it stands (a few minutes)
Run python3 demo.py first and write the numbers down. That is your before number.
When the provider does throttle us, it says how long to wait: a RateLimited carries a
retry_after. app/drafting.py ignores it. It retries immediately, as fast as the loop goes
round, and when it finally gives up it tells the caller to come back in exactly one second no
matter what the provider said. So ten throttled workers spend their attempts in a few
milliseconds, all at the same instant, and the caller above them is given a number that was never
true.
Make the retry path wait, and make the number it hands back real. Waiting the same amount on
every worker is its own problem - if they were all throttled together they will all come back
together. A one-line comment saying why you chose what you chose is enough.
2. Give the fleet one shared budget (the main task)
This is the real task. The limit has to hold across all ten workers, not inside each one.
shared_store.py is the store every worker can reach. Read its docstring before you start,
particularly run_atomic. Reading a value, deciding from it and writing it back as three separate
calls is not safe here: between your read and your write another worker reads the same value and
you both think there is room. run_atomic is the way to read, decide and write as one step.
Two things have to be counted, not one:
- Requests. 60 a minute across the fleet.
- Tokens. 20,000 a minute across the fleet. This is the one that catches people out. A
classify call costs a couple of hundred tokens and a draft_section call costs several
thousand, so counting requests tells you almost nothing about whether you are near the token
ceiling. You have to decide what a call will cost before you make it, and what to do when it
turns out to cost something different. provider.estimate_prompt_tokens(prompt) is there to
help with the first half.
Both budgets have to be taken together. If you spend a request and then find there are no tokens
left, you have already spent the request, and nothing gives it back.
What has to keep working, because the rest of the product calls it:
| what | contract |
|---|---|
| DraftingClient(provider, store=store) | keeps working, with the store as a keyword argument |
| call_model(prompt, max_output_tokens=512) | returns a provider.Response, or raises provider.RateLimited when it could not get through |
| summarise(text), classify(text), draft_section(matter_id, section, notes) | same names, same arguments, same return types |
limiter.py can stay, change or go. It is a reasonable local guard, but it cannot be the thing
that enforces a fleet-wide number.
You will know it is working when demo.py shows the provider rejecting almost nothing. Right now
it rejects hundreds.
3. Stretch: decide what happens when the store is not there
The shared store is now the thing standing between us and the provider's quota, which means it is
also a new way for drafting to break. store.fail_all(True) makes every call raise
StoreUnavailable.
There is no free answer here. Refusing every request protects the quota and the bill and takes
drafting down. Letting every request through keeps drafting up and puts the account over its
limit, which is how you get throttled for the rest of the hour. Pick one, make it a constructor
argument called fail_open that defaults to whichever you think is right for this product, and
say in your notes why.
If you still have time: give each law firm its own slice of the budget with a tenant argument,
so one firm running a bulk import cannot starve everyone else.
Start here
python3 demo.py - write down the 429 count.
python3 -m unittest discover -s tests - they pass.
- Read
app/drafting.py, about 60 lines. Phase 1 and phase 2 both happen there.
- Read the
shared_store.py docstring, particularly run_atomic.
What's here
app/drafting.py the worker: builds a limiter, calls the model, retries
limiter.py the per-process token bucket it builds
shared_store.py the store every worker can reach, with run_atomic
provider.py the model provider stand-in: enforces the quota, reports token cost
demo.py runs a fleet of workers and prints what the provider thought of it
tests/ tests for the stand-ins and for one worker (they pass)
Writing it down
When you are done, a short NOTES.md, three or four lines: your before and after numbers, what
you counted and why, and what you would want to watch in production once this shipped.
Running it
python3 demo.py # the fleet, and the provider's verdict
python3 demo.py --workers 20 --jobs 8 # a different shape of load
python3 -m unittest discover -s tests -v # the tests
Python 3.11, standard library only. Nothing to install, no network access needed.
Time
About 30 minutes of work: roughly 5 minutes on phase 1 and 25 on phase 2. Phase 3 is there for
people who get through the first two with time to spare; most will not, and that is the expected
outcome. If you run out of time mid-way, say in a sentence or two what you would have done next.