The assistant cites the wrong passage
Context
Our assistant answers questions about patents and office actions, and it keeps pointing at the wrong part of the document. It searches a set of patent documents, takes the top 3 passages, puts them in a prompt, and asks a model to answer with a citation. Attorneys report answers that cite the claims when they asked about the background, or that quote half a sentence. The model here is a simulated stand-in, so results are the same on every run.
Try it:
python3 -m rag.answer "What coolant flow rate does claim 1 of the coolant manifold patent require?"
Your task
eval_retrieval.py is already written, so you can see the problem in numbers
before you touch anything.
- Run
python3 eval_retrieval.py and write the numbers down. It prints hit@1
and hit@3, and then the questions that miss with what came back instead, which
is usually the fastest way in. hit@3 is how often the right passage is in the
top 3, hit@1 how often it is the first result.
- Fix the two things that are wrong. Chunking cuts the text every 200
characters regardless of what the text says, so chunks begin and end mid
sentence and a heading gets separated from the section it belongs to. And
scoring counts query words, so a chunk full of common words wins. Keep
build_index() and index.search(question, k) returning (score, chunk)
pairs so rag/answer.py keeps working.
- Re-run the harness and write a short
NOTES.md, three or four lines: what was
wrong, what you changed, and hit@1 and hit@3 before and after.
You do not need anything clever for the scoring. Weighting a word by how rare it
is across the corpus is enough to move the numbers.
Start here
rag/chunker.py splits documents into chunks. Look at what it does to a heading.
rag/index.py scores one chunk against a question. This file is very short.
rag/answer.py retrieves the top 3 and builds the citation. Leave it alone.
What's here
rag/chunker.py splits documents into fixed 200-character chunks
rag/index.py scores chunks by counting query words
rag/answer.py retrieves top 3 chunks, prompts the (simulated) model, returns answer + citations
data/patents/*.md 8 patent documents with "#" titles and "##" sections
eval/questions.json 12 questions with the expected doc and section
tests/test_rag.py current tests (all passing)
Each document is one patent publication in plain text: title, abstract, technical field, background, summary, drawings, detailed description and claims.
Running it
python3 -m unittest discover -s tests -v
python3 -m rag.answer "What does the background of the coronary stent patent say about durable polymer coatings?"
python3 eval_retrieval.py # the eval harness, works as shipped
Python 3.11, standard library only. No installs needed.
Time
Aim for about 30 minutes. You do not need to finish everything. We care more
about how you approach it than about how much you get through.