Which office action items need an attorney?
Context
From one of our patent attorneys, last week: "The queue told me claim 1 of the
crystallizer case was limited to 6.0-125.0 mm. The office action says 12.5. I
spent an hour drafting a response to a number the model made up, and it was
sitting at the bottom of my list while three dependent claims sat at the top."
We read office actions with a model, turn each one into a list of items to work
on, and hand that list to the docketing team, and right now we pass on whatever
the model says in whatever order it comes out.
Your task
1. The queue is in the wrong order. review_document() promises the items
come back most urgent first, and they don't. Find it and fix it. This is a
warm-up; it should take a few minutes.
2. Decide which items an attorney has to look at. Today nothing is checked:
needs_attorney is always empty, so everything the model says is taken on
trust. Add the check. For every item, the numbers and the patent numbers the
model reported must actually appear in the document text; if they don't, an
attorney has to see that item.
Harmless differences in how a value is written must pass:
| the document says | the model reports | verdict |
| --- | --- | --- |
| between 5 and 15 wt% | 5.0, 15 | fine |
| between 1,200 and 1 500 L/h | 1200.0, 1500.0 | fine |
| US 2019/0555111 A1 | US20190555111A1 | fine |
| EP 3 100 964 A1 | EP3100964 (no kind code) | fine |
| between 3.3 and 9.0 rev/s | 33.0 | needs an attorney |
| between 1,200 and 4,000 L/h | 120 | needs an attorney |
| US 2019/0555111 A1 | US20190555112A1 | needs an attorney |
One changed digit, one extra digit, or a decimal point in the wrong place is
exactly what a made-up value looks like, so those must be caught. Nothing is
dropped or corrected on the way: an item stays in items with the values the
model gave it, it just gets flagged.
review_document(text, client=None) returns exactly three keys, and the tests
we run against your work read only these:
{
"publication_number": str | None, # whatever the model called the document
"items": [item, ...], # every item, most urgent first
"needs_attorney": [item_id, ...], # the item_ids an attorney has to see,
# in the same order they appear in items
}
An item keeps the keys it has now (item_id, kind, urgency, summary,
values, references) — see build_items(). Add extra keys to an item if they
help a reviewer; nothing we run will mind.
3. Stretch — one bad reply must not sink the batch. review_batch() runs a
whole day's post through the model. Sometimes the model answers with an
apology, or with truncated JSON, instead of the object we asked for. That one
document should come back as something an attorney has to look at, and the rest
of the run should finish normally.
Finishing all three is not expected. Phase 1 and phase 2 are the job. Phase
3 is there in case you have time left, and most people will not reach it.
Start here
extract.py is the whole pipeline: it asks the model, turns the answer into
items, orders them, and returns the result. Start by running it over the
documents in data/actions/ and reading the output next to the documents
themselves — two of the six have something wrong with them. Then read
review_document() and order_items().
llm.py is the model. Treat it as a service someone else runs: don't edit it,
and don't build anything that depends on which documents it happens to get
wrong.
Put your checking code wherever you like — a new module is fine. Add tests for
it in tests/. If you make a judgement call about what counts as the same
value, leave a line about it in a NOTES.md or a comment.
What's here
extract.py the pipeline: prompt, model call, items, ordering, result
llm.py the model client (simulated, deterministic, offline). Don't edit.
data/actions/ six office actions as text, each with a front page and its claims
tests/ the tests we have now. They pass; keep them passing.
Running it
python3 extract.py # review every document and print the queue
python3 extract.py data/actions/us9871234.txt # just one
python3 -m unittest discover -s tests -v # run the tests
Python 3.11, standard library only, no network.
Time
About 30 minutes. Phase 1 should take a few minutes and phase 2 is the bulk of
it. We are more interested in how you decide what counts as the same value, and
in what you choose to test, than in how much you get through.