Why did this patent come up?
Context
Our prior-art search tool takes what an attorney types and hands back a ranked list
of patents, each with an id, a title and a relevance score. One attorney put the
complaint plainly: "You have told me these patents are relevant. You have not told
me why, so I still have to open every one and read it." The ranking itself is
in decent shape — the right patent is almost always first — but the results page
gives no reason for its answer, and the pieces needed to give one are already in
the code.
Your task
Three phases, in order. Phase 3 is the stretch — see the Time section below.
Phase 1 (warm-up, about 5 minutes)
Ask for more than ten results and you still get ten: index.search(query, 30) hands
back ten chunks however many matched. search_patents builds its results out of those
chunks, so patents that belong in the list quietly go missing — compare
search_patents("silver coating percent by weight", limit=6) against the six
documents in data/patents/. Find it, fix it, and make sure asking for a small
number still gives a small number.
Phase 2 (the main task, about 25 minutes)
Every result should carry the passage that made the patent come up, and say where in
the document that passage lives.
- Each result gains two fields:
passage, the one or two sentences from that patent
that best match what was searched for, and section, the heading of the ##
section the passage came from ("Claims", "Background", and so on).
- The existing
patent_id, title and score stay as they are, and one result per
patent stays one result per patent.
- The passage is quoted from the document, word for word. No summarising, no cutting
a sentence in half, and no heading lines mixed into it.
- The passage should follow the search: someone searching claim language should get
claim text, someone asking what the background says should get the background.
rag/chunker.py already splits each document into whole-sentence chunks that carry
their heading, and rag/index.py already scores those chunks. This phase is about
wiring what is there into the result, not about writing a chunker or a new scorer.
eval_search.py measures three things: how often the right patent comes back first,
how often it is in the top three, and how often the passage shown comes from the
section an attorney said was the right one. The first two are already good and should
stay that way. The third is what this phase moves.
Phase 3 (stretch — most people will not reach this)
Sometimes a patent is only in the list because a handful of ordinary words are
scattered around it, and nothing in it really matches. Once phase 2 is in, such a
patent still gets a quoted passage — the best of a bad set, an opening line or some
unrelated detail — and on the page that reads as if it were the reason for the hit.
Make the result say there is no clearly matching passage instead: passage and
section are None, and the printed output says so.
A rule of thumb that works here: if the best passage in a patent covers less than
about half of the meaningful words that were searched for, there is no real match.
Whatever rule you pick, a patent that genuinely matches must not be written off.
Start here
rag/search.py — rolls chunk hits up into one result per patent. This is where all
three phases land. Note what it does with each chunk it has just been handed.
rag/chunker.py — already produces whole-sentence chunks that know their section.
eval_search.py — run this before you change anything, and again afterwards.
What's here
rag/chunker.py splits a patent into whole-sentence chunks that keep their heading
rag/index.py scores chunks against a search, best first
rag/search.py rolls chunk hits up into one result per patent (the results page)
eval_search.py 12 reviewed searches with the patent and section an attorney expected
data/patents/*.md 6 patent documents, "#" title and "##" sections
tests/test_search.py the current tests, all passing
Each document is one patent publication in plain text: title, abstract, technical
field, background, summary, drawings, detailed description and claims.
Running it
python3 -m unittest discover -s tests -v
python3 -m rag.search "liquid cooling channels for prismatic battery cells"
python3 eval_search.py
python3 eval_search.py -v # also lists the searches that came back wrong
Python 3.11, standard library only. No installs, no network.
Time
Aim for about 30 minutes. Finishing all three phases is not expected, and most
people will not reach phase 3 at all — phase 1 and phase 2 are the substance of the
exercise. We care more about how you work than about how much you finish, so leave
the code in a state you would be happy to hand to a colleague and say out loud what
you would do next.