I'd built a small local RAG server โ PDFs and notes chunked, embedded with sentence-transformers, indexed in FAISS, served over MCP so an LLM (or a CLI, or a browser) could search them semantically. It worked. Scores looked healthy, 0.5โ0.6 cosine similarity, top results from the right documents.
Then I actually read what it returned.
A confident answer, built from nothing
Querying "attention mechanism in transformers" against the Attention Is All You Need paper, result #2 came back scored 0.57 โ respectable โ with this as its supporting text:
ng process more difficult . <EOS> <pad> <pad> <pad> <pad> <pad> <pad> Figure 3: An example of the attention mechanism following long-distance dependencies in the encoder self-attention...
That's not a passage about attention. It's the tail end of a figure caption, sliced mid-sentence, with stray decoder padding tokens still attached. The score said trust this. The text said this makes no sense. One of them was lying.
The chunk text was never stored
The ingestion pipeline chunks a document, embeds each chunk, and indexes the vector in FAISS โ but the chunk-to-metadata table only recorded doc_id, chunk_index, and page. The actual text of the chunk was embedded, then thrown away.
So at search time, search_knowledge had to guess where the matched chunk had been, after the fact:
# chunks are ~500 chars, sentence-boundary aware, 50-char overlap โ # none of that is uniform, but this pretends it is: chars_per_chunk = len(full_text) // doc_meta.chunk_count start = r["chunk_index"] * chars_per_chunk end = start + chars_per_chunk + 100 chunk_text = full_text[start:end].strip()[:600]
Real chunks are variable-length, sentence-aware, and overlapping. This math assumes they're all identical. The error compounds with every chunk index, so a match deep in a 97-chunk document could land anywhere near it โ including inside a completely unrelated figure caption. The bug wasn't in the embeddings or the ranking. It was in the one step that should have needed no cleverness at all: remembering what you already had.
Stop reconstructing. Start storing.
The fix removes a heuristic instead of adding one โ store the real chunk text once, at ingestion, and hand it back verbatim.
+ chunk_text TEXT NOT NULL DEFAULT '' -- added to chunk_mapping
chunk_text = r["text"] # the real thing, not a guess
~15 lines of reconstruction logic deleted. I rebuilt the local knowledge base from source and confirmed directly against SQLite: 1,587 / 1,587 stored chunks, zero empty. The same query now returns a clean, complete sentence โ because it's no longer reconstructing one.
Stop eyeballing it. Measure it.
A fixed bug is a relief. It isn't evidence. So before touching anything else, I built an eval harness: 24 hand-labeled (query โ expected document) pairs across all 6 indexed documents, scored on Recall@1, Recall@5, and MRR โ the same metrics an IR team would report, run against the same code path production search uses.
Baseline (dense embeddings only):
| Mode | Recall@1 | Recall@5 | MRR |
|---|---|---|---|
| Dense (FAISS cosine) | 95.8% | 100% | 0.979 |
One miss, out of 24: "few-shot in-context learning with intermediate reasoning steps" ranked the LoRA paper first, when the correct answer was Chain-of-Thought. rank 2
Hybrid search โ a tie, not a win
The obvious next move: add BM25 keyword search alongside the dense embeddings, and fuse the two rankings with Reciprocal Rank Fusion. Exact terms and paraphrases both get a vote.
| Mode | Recall@1 | Recall@5 | MRR |
|---|---|---|---|
| Dense only | 95.8% | 100% | 0.979 |
| BM25 only | 91.7% | 100% | 0.958 |
| Hybrid (RRF) | 95.8% | 100% | 0.979 |
Hybrid tied dense-only. Not a regression โ but not a win either, and I'd rather report that plainly than dress it up. The reason became obvious on inspection: on the one query that still missed, dense and BM25 independently agreed on the wrong document. Fusing two rankers that already agree with each other โ just not with the truth โ can't out-vote a shared mistake. RRF only helps when the signals disagree.
A reranker that owes neither signal anything
A cross-encoder doesn't combine two independent similarity scores โ it reads the query and the candidate passage together, in one forward pass, and scores the pair directly. It isn't built from dense or BM25's opinion, so it isn't stuck agreeing with either. I added one (ms-marco-MiniLM-L-6-v2) as a final stage over the fused candidate pool.
| Mode | Recall@1 | Recall@5 | MRR |
|---|---|---|---|
| Dense only | 95.8% | 100% | 0.979 |
| BM25 only | 91.7% | 100% | 0.958 |
| Hybrid (RRF) | 95.8% | 100% | 0.979 |
| + Cross-encoder rerank | 100% | 100% | 1.000 |
The last miss resolved. rank 1 Full marks across all 24 queries.
Two lessons, not one
Storing beats reconstructing. The original bug wasn't a hard problem solved poorly โ it was a missing column standing in for an unnecessary heuristic. When a system starts approximating something it has already computed once, that's usually a sign the value should have been saved the first time.
An eval harness turns "it works" into a real claim. Without one, hybrid search would have shipped as a plausible improvement โ it never regressed anything, after all. With one, the honest result was visible immediately: a tie, for a diagnosable reason, pointing straight at what would actually fix it. Measuring first is what made the second fix targeted instead of speculative.
- FAISS
- sentence-transformers
- rank-bm25
- cross-encoder/ms-marco-MiniLM-L-6-v2
- SQLite
- MCP