Run drafting on a cheaper model, keep the current one as backup
Context
Our model bill has been growing faster than the number of patents we process, and finance has
asked us to stop paying premium per-token rates for two small jobs: the short summary we show
under each patent document, and the CPC subclass codes we suggest for a draft. We now have a
self-hosted open-source model running in our own cluster that is good enough for both, and it
costs us nothing per call. We want to be able to move those two jobs onto it, and move them
back, without editing application code.
Both models in this repo are deterministic offline stand-ins. They behave like the real ones,
including their differences: the self-hosted model writes shorter summaries and recognises fewer
technical synonyms, so it will not always agree with the hosted one.
Your task
Three phases, in order. Phase 3 is a stretch goal. You are not expected to finish all three,
and most people do not reach phase 3 - we care much more about how you work through phases 1
and 2 than about getting to the end.
1. Fix a bug in the code as it stands (a few minutes)
The hosted model tells us when it stopped a reply early because it ran out of room (a cut-off,
incomplete answer). One of the two features handles that badly: when it happens, the caller is
told something that is not true, and nothing anywhere records that a call failed. Find it, decide
what the right behaviour is, and change it. A one-line note in your commit or a comment saying
why you chose that behaviour is enough.
2. Make the provider a configuration choice, not a code change (the main task)
Right now app/summarize.py and app/classify.py each build a request for the hosted provider
by hand: they import its client, name its model (atlas-large-2), and use its request fields
(input, max_output_tokens, stop_reason). Both features ask a model the same simple thing -
"here is a prompt and a system instruction, give me back text" - but each one says it in the
hosted provider's private vocabulary.
Introduce one interface that both features use: a single small thing, written once, that
takes a prompt (plus a system instruction, a token limit and a temperature) and returns the text
of the reply. Both features call that, and only that. Behind it, write the code that actually
talks to a provider: one version for the hosted client we use today, and one version for a server
that speaks the common /v1/chat/completions HTTP shape (the self-hosted model, and the way vLLM
and Ollama serve models). Which of the two is used is decided by configuration, at the moment a
request is made:
| variable | meaning |
|-------------------------|---------------------------------------------------------------------|
| MODEL_PROVIDER | hosted or local. If it is not set, keep using hosted - deploying your change must not move traffic on its own. |
| LOCAL_MODEL_URL | base URL of the self-hosted server, e.g. http://127.0.0.1:8089/v1 |
| LOCAL_MODEL_NAME | optional, model name to send to it (default open-llm-8b-instruct) |
| MODEL_TIMEOUT_SECONDS | optional, how long to wait for the self-hosted model (default 5) |
Read those variables when a request is made rather than once at import, so flipping the flag does
not need a restart.
Two constraints:
- Put the interface and the provider-specific code outside
app/ - a new top-level module is
fine. When you are done, no file under app/ should name a provider, import provider code, or
read those environment variables. app/ should not be able to tell which model answered.
app.summarize.summarize(text, max_words=40) and app.classify.suggest_codes(text) keep their
names, their arguments and their return types. Callers elsewhere in the product are not in this
repo, and we are not changing them.
3. Stretch: fall back to the hosted model when the self-hosted one lets us down
The self-hosted cluster is ours, which means it is sometimes slow, restarting or overloaded. With
MODEL_PROVIDER=local, a request that takes too long or fails should quietly be served by the
hosted model instead, so users see an answer. Not every failure deserves that: think about which
ones mean "this box is having a bad minute" and which ones mean "this request is wrong and the
hosted model would reject it too" - paying twice to get the same rejection is worse than failing.
Start here
- Run the tests:
python3 -m unittest discover -s tests (they pass).
- Read
app/summarize.py and app/classify.py - they are about 25 lines each. That is where
phase 1 and phase 2 happen.
- Skim
hosted_client.py (the paid provider we use today) and local_model_server.py (the
stand-in for our own cluster) to see how differently the two speak.
- Add tests for what you change. The existing tests must keep passing.
What's here
app/summarize.py short summary of a patent document; calls the hosted provider directly
app/classify.py CPC subclass code suggestions; calls the hosted provider directly
hosted_client.py client for the paid hosted provider, in its own request/response shape
local_model_server.py stand-in for our self-hosted model, speaking /v1/chat/completions
simulated_models.py the deterministic behaviour behind both stand-ins (no need to edit)
tests/test_app.py tests for the two features
tests/test_stand_ins.py tests for the two stand-ins
The self-hosted stand-in has the rough edges of a real deployment: the first request is slow while
the model warms up, it can return 503 when overloaded, and it has a small context window. In a
test you can start one on a free port:
import local_model_server
server = local_model_server.start_in_thread(latency=0.0, first_request_latency=0.0)
print(server.url)
print(server.request_count)
local_model_server.stop(server)
Options: latency, first_request_latency, fail_every (503 on every Nth request),
force_status (always reply with this HTTP status), max_context_chars.
Running it
python3 -m unittest discover -s tests # the test suite
python3 local_model_server.py --port 8089 # run the self-hosted stand-in by hand
# once phase 2 is done, this should go to the stand-in above instead of the paid provider
MODEL_PROVIDER=local LOCAL_MODEL_URL=http://127.0.0.1:8089/v1 python3 -c \
"from app.summarize import summarize; print(summarize('A cell stack includes a lithium anode. It resists dendrites.'))"
Python 3.11, standard library only. Nothing to install, no network access needed.
Time
About 30 minutes of work: roughly 5 minutes on phase 1 and 25 on phase 2. Phase 3 is there for
people who get through the first two with time to spare; most will not, and that is the expected
outcome. If you run out of time mid-way, say in a sentence or two what you would have done next.