Show the draft as the model writes it
Context
An attorney asks for an abstract and then watches a spinner. "The model writes about
four hundred words and I get nothing at all until it has written every one of them —
and I cannot tell whether it is thinking or whether it has died." Today the editor
makes one request, waits for the whole draft to come back in one lump, and only then
puts anything on the screen.
Your task
Work in whatever order you like, but read the numbers from node measure.js first.
1. A bug to warm up on (about 5 minutes). When a request for a draft fails, the
editor takes the spinner down and shows the attorney a finished draft that happens to
be empty, instead of saying something went wrong. The last two runs of
node measure.js both show it. It is a small mistake somewhere between the request
going out and the state the editor renders. Fix it.
2. Make the draft appear as it is written (the main task, about 25 minutes). Text
should reach the browser as the model produces it, and the editor should grow a piece
at a time instead of appearing all at once at the end. This needs a change on the
Python side and a change on the browser side: the server has to send each piece as
soon as it has it rather than collecting them all first, and the client has to read
the response as it arrives rather than waiting for the end of it.
The two sides already agree on a format for this, and encode_event in
stream_server.py writes it. One piece of draft text looks like this on the wire and
ends with a blank line:
id: 12
event: token
data: electrode
There is exactly one space after data:, and a reader takes that one space off and
keeps everything after it, so a piece that begins with a space (" electrode")
survives the trip. The response is sent with the headers in STREAM_HEADERS: the
content type is text/event-stream, and there is deliberately no Content-Length,
because we do not know how long the draft will be and a length would make the browser
wait for all of it.
Two things to watch:
- The finished text must be byte for byte what the model produced. Pieces are not
words. A piece can be a single space, or the first four letters of a long word, so
nothing between the model and the editor may trim a piece, drop an empty-looking
one, or put the text back together on word boundaries. Getting this wrong shows up
as
patentapplication and as a draft that is a few characters short.
- A read from the network gives whatever bytes happened to arrive. One read can
hold several events, or half an event, or the tail of one event and the head of the
next. Whatever the client does, an event split across two reads must not be lost,
cut short, or applied twice.
3. Stretch — most people will not reach this, and that is expected. A connection
that drops part way through a draft currently ends up looking exactly like a finished
draft, because by the time the model dies the response headers have long gone and
there is no status code left to change. Make the two cases different. A draft that
finished should end with one last event, event: done, and the client should treat a
response that ends without that marker as an error rather than as a finished draft: a
done marker puts the state in status: 'complete', and a response that ends any other
way must come out as something other than 'complete', must carry an error the
editor can show, and must keep the text that did arrive.
If you have time at the end, a short NOTES.md (5-10 lines) helps us: what you
changed, the before and after numbers from node measure.js, and what you would do
next.
Names the rest of the app already uses
Keep these working, in these places, with these names.
stream_server.py: create_app(options) with its prompts, model and
delay_ms options, the route GET /api/drafts/<id>/stream, DEFAULT_PROMPTS,
encode_event(event_id, name, data), and the test helper
read_response(app, draft_id, headers=None, query=""), which returns
(status, body) where body is whatever the application returned.
store.py: save_draft(draft_id, text) and get_draft_text(draft_id). A draft
that finished is stored whole, as it is now.
model_stream.py: stream_tokens(prompt, options). This stands in for the real
model. Don't edit it.
web/stream_client.js: createState(), applyChunk(state, event),
finishStream(state) and streamDraft(draftId, onState, options) with its
fetchImpl and baseUrl options. streamDraft calls onState(state) every time
the state changes, which is how the editor knows to re-render.
- The state object the editor renders:
text (the draft so far, exactly as it should
appear), tokenCount (how many pieces of text have been folded in), status and
error.
Start here
node measure.js
It starts the server on a spare port, drives web/stream_client.js against it the way
the editor does, and prints three runs: a whole draft, a draft where the model
connection drops part way through, and a draft id that does not exist. For each one,
write down how many times the editor was updated, how many bytes of the response had
arrived when the first character showed up, how long the attorney had waited by then,
and the status the editor was left in. Those are the numbers your change has to move.
Then run the tests, so you know they pass before you start.
What's here
stream_server.py the endpoint (a WSGI app), and the wire format the front end expects
model_stream.py drafting model client (simulated). Treat it as external: don't edit it.
store.py where a finished draft is kept
web/stream_client.js browser side: make the request, fold the text into editor state
measure.js starts the server and drives the client, printing what the editor got
data/drafts.json the text the simulated model draws its drafts from
tests/ existing tests. They pass now; keep them passing.
model_stream.py is deterministic: the same prompt always produces the same pieces,
so you never need a network or a real model. The pieces do not line up with words.
Passing it {"disconnect_after": 12} makes the upstream connection drop after twelve
pieces, which is how both the good path and the bad path get tested.
Running it
node measure.js # what the editor gets, and when
python3 -m unittest discover -s tests -v # python tests
node --test --no-warnings $(ls tests/*.test.js) # browser tests (Node 22)
python3 stream_server.py serve # serve on http://127.0.0.1:8000
curl -N http://127.0.0.1:8000/api/drafts/d-1001/stream # watch the raw response
DRAFT_DISCONNECT_AFTER=40 python3 stream_server.py serve # make the model drop part way
DRAFT_DELAY_MS=0 python3 stream_server.py serve # no pause between pieces
curl -N tells curl not to hold the response back, so you see the bytes as the server
sends them. The server waits 5 ms after each piece by default, so a local run feels
like a real model instead of finishing instantly.
Python 3.11, standard library only. Node 22. No npm install, no pip install.
Time
About 35 minutes. Phase 1 is a five minute warm-up, phase 2 is the task we care about,
and phase 3 is a stretch. You are not expected to finish all three phases — most
people do not, and that is fine. We are more interested in how you work out what is
happening and what you decide to do about it than in how much of the list you get
through.