The right patent doesn't come up
Context
Searchers keep telling us that the patent they know exists is not in the first
ten results. The same patent turns up fine when they try either search method on
its own, so the problem is in the list we actually show them, which is built by
merging the two. Patent search here runs two methods over the same records and
merges their results: BM25 (a keyword scoring method: it rewards rare words that
match, gives less credit for repeats, and adjusts for document length) and a
meaning based method built on embeddings (an embedding is a list of numbers that
stands for the meaning of a piece of text, so that two ways of saying the same
thing get two similar lists of numbers). Nobody has measured the merge since it
was first put together.
Your task
- Write
eval_search.py at the repo root and get the starting numbers before
you change any search code. It reads eval/queries.json, runs each query
through hybrid_search, and prints three numbers over the whole query set:
recall@1 (how often the right patent is the very first result), recall@10 (how
often the right patent shows up anywhere in the top 10) and MRR (mean
reciprocal rank: for one query, take the position of the first right answer
and use 1 divided by that position, so 1.0 at rank 1, 0.5 at rank 2, 0.1 at
rank 10, and 0 if it never shows up at all, then average that over the
queries). Expose a function recall_at_k(retrieved_ids, expected_ids, k) so
that other code can import it. Write the three starting numbers down before
you touch anything else.
- Find and fix what is wrong in
search/. The merge step is the first place to
look, and it is not the only thing wrong. Treat search/embedder.py as a
fixed outside service: you can change what text you send it and what you do
with the numbers it returns, but not the numbers themselves.
- Keep these names where they are, because other code imports them:
BM25Index in search/bm25.py, DenseIndex in search/dense.py,
embed(text) in search/embedder.py, and
hybrid_search(query, k=10) in search/hybrid.py, which returns a ranked
list of (score, doc) pairs, best first. Keep hybrid_search getting its two
candidate lists from BM25Index.search and DenseIndex.search.
- Re-run the harness and record the after numbers next to the before numbers.
If a change makes one number better and another worse, say so.
- Write a short
NOTES.md (5 to 10 lines): what you found, what you changed,
what you traded off, and the before and after numbers.
Do not edit data/patents.json or eval/queries.json. If you run out of time, a
clear note about what you would do next is worth more than a half finished change.
What's here
search/corpus.py loads data/patents.json and builds the text each method reads
search/bm25.py BM25Index: the keyword method
search/embedder.py embed(text): the simulated embedder, 48 numbers per text
search/dense.py DenseIndex: the meaning based method, ranked by cosine
similarity (how closely two embeddings point the same way)
search/hybrid.py hybrid_search(query, k): the merge step, what /search calls
eval/queries.json 15 queries, each naming the document ids that should come back
data/patents.json 43 patent records: id, publication number, filing date,
title, abstract, one representative claim, CPC codes (the
subject labels an examiner attached to the document)
data/concepts.json the topic word lists the simulated embedder is built from
tests/ tests that describe how search behaves today
There is no network in this repo. search/embedder.py is a deterministic
simulated stand-in for the hosted embedding service, not a call to it. It behaves
like a small real one in the ways that matter here: the same text always gives
the same numbers, two wordings of one idea land close together, unrelated text
lands far apart. It is also wrong in the ways a real one is wrong: it ignores
word order, it only knows the topics in data/concepts.json, and it waters down
long documents. data/patents.json is a small synthetic stand-in with the same
row shape as the production table. CPC is the patent classification system, so
"G01S 17/89" is one of those subject labels, and searchers do paste codes and
publication numbers straight into the box.
Running it
From the repo root:
python3 -m unittest discover -s tests -v # 37 tests, all passing on the starter
python3 eval_search.py # once you have written it
To look at one query while you work:
python3 -c "
from search.hybrid import hybrid_search
for score, doc in hybrid_search('cells of an electric car getting too hot', 10):
print(round(score, 3), doc['id'], doc['title'])
"
To see what each method says on its own:
python3 -c "
from search.bm25 import BM25Index
from search.dense import DenseIndex
from search.corpus import load_corpus
docs = load_corpus()
query = 'cells of an electric car getting too hot'
print('keyword:', [(round(s, 2), d['id']) for s, d in BM25Index(docs).search(query, 5)])
print('meaning:', [(round(s, 2), d['id']) for s, d in DenseIndex(docs).search(query, 5)])
"
Everything runs in well under a second.
Time
Aim for about 25-30 minutes. You don't need to finish everything; we care more
about how you approach it than about completeness.