CASE knowledge-mcp-server FILED 2026-08-15 SUBSYSTEM search_knowledge / EmbeddingStore

The Ghost Chunk

A retrieval bug that returned confident, well-scored answers built from text that was never actually retrieved โ€” and the three-stage fix that ended in a perfect eval score.

Setup

I'd built a small local RAG server โ€” PDFs and notes chunked, embedded with sentence-transformers, indexed in FAISS, served over MCP so an LLM (or a CLI, or a browser) could search them semantically. It worked. Scores looked healthy, 0.5โ€“0.6 cosine similarity, top results from the right documents.

Then I actually read what it returned.

The Symptom

A confident answer, built from nothing

Querying "attention mechanism in transformers" against the Attention Is All You Need paper, result #2 came back scored 0.57 โ€” respectable โ€” with this as its supporting text:

search_knowledge โ€” result #2, page 13
ng
process
more
difficult
.
<EOS>
<pad>
<pad>
<pad>
<pad>
<pad>
<pad>
Figure 3: An example of the attention mechanism following
long-distance dependencies in the encoder self-attention...

That's not a passage about attention. It's the tail end of a figure caption, sliced mid-sentence, with stray decoder padding tokens still attached. The score said trust this. The text said this makes no sense. One of them was lying.

Root Cause

The chunk text was never stored

The ingestion pipeline chunks a document, embeds each chunk, and indexes the vector in FAISS โ€” but the chunk-to-metadata table only recorded doc_id, chunk_index, and page. The actual text of the chunk was embedded, then thrown away.

So at search time, search_knowledge had to guess where the matched chunk had been, after the fact:

src/knowledge_mcp/server.py โ€” before
# chunks are ~500 chars, sentence-boundary aware, 50-char overlap โ€”
# none of that is uniform, but this pretends it is:
chars_per_chunk = len(full_text) // doc_meta.chunk_count
start = r["chunk_index"] * chars_per_chunk
end = start + chars_per_chunk + 100
chunk_text = full_text[start:end].strip()[:600]

Real chunks are variable-length, sentence-aware, and overlapping. This math assumes they're all identical. The error compounds with every chunk index, so a match deep in a 97-chunk document could land anywhere near it โ€” including inside a completely unrelated figure caption. The bug wasn't in the embeddings or the ranking. It was in the one step that should have needed no cleverness at all: remembering what you already had.

The Fix

Stop reconstructing. Start storing.

The fix removes a heuristic instead of adding one โ€” store the real chunk text once, at ingestion, and hand it back verbatim.

src/knowledge_mcp/storage/metadata.py
+ chunk_text TEXT NOT NULL DEFAULT ''  -- added to chunk_mapping
src/knowledge_mcp/server.py โ€” after
chunk_text = r["text"]  # the real thing, not a guess

~15 lines of reconstruction logic deleted. I rebuilt the local knowledge base from source and confirmed directly against SQLite: 1,587 / 1,587 stored chunks, zero empty. The same query now returns a clean, complete sentence โ€” because it's no longer reconstructing one.


Method

Stop eyeballing it. Measure it.

A fixed bug is a relief. It isn't evidence. So before touching anything else, I built an eval harness: 24 hand-labeled (query โ†’ expected document) pairs across all 6 indexed documents, scored on Recall@1, Recall@5, and MRR โ€” the same metrics an IR team would report, run against the same code path production search uses.

Baseline (dense embeddings only):

ModeRecall@1Recall@5MRR
Dense (FAISS cosine)95.8%100%0.979

One miss, out of 24: "few-shot in-context learning with intermediate reasoning steps" ranked the LoRA paper first, when the correct answer was Chain-of-Thought. rank 2

Attempt 1

Hybrid search โ€” a tie, not a win

The obvious next move: add BM25 keyword search alongside the dense embeddings, and fuse the two rankings with Reciprocal Rank Fusion. Exact terms and paraphrases both get a vote.

ModeRecall@1Recall@5MRR
Dense only95.8%100%0.979
BM25 only91.7%100%0.958
Hybrid (RRF)95.8%100%0.979

Hybrid tied dense-only. Not a regression โ€” but not a win either, and I'd rather report that plainly than dress it up. The reason became obvious on inspection: on the one query that still missed, dense and BM25 independently agreed on the wrong document. Fusing two rankers that already agree with each other โ€” just not with the truth โ€” can't out-vote a shared mistake. RRF only helps when the signals disagree.

Attempt 2

A reranker that owes neither signal anything

A cross-encoder doesn't combine two independent similarity scores โ€” it reads the query and the candidate passage together, in one forward pass, and scores the pair directly. It isn't built from dense or BM25's opinion, so it isn't stuck agreeing with either. I added one (ms-marco-MiniLM-L-6-v2) as a final stage over the fused candidate pool.

query dense (FAISS) BM25 RRF fuse candidatepool (15) re-rank
ModeRecall@1Recall@5MRR
Dense only95.8%100%0.979
BM25 only91.7%100%0.958
Hybrid (RRF)95.8%100%0.979
+ Cross-encoder rerank100%100%1.000

The last miss resolved. rank 1 Full marks across all 24 queries.


Takeaways

Two lessons, not one

Storing beats reconstructing. The original bug wasn't a hard problem solved poorly โ€” it was a missing column standing in for an unnecessary heuristic. When a system starts approximating something it has already computed once, that's usually a sign the value should have been saved the first time.

An eval harness turns "it works" into a real claim. Without one, hybrid search would have shipped as a plausible improvement โ€” it never regressed anything, after all. With one, the honest result was visible immediately: a tie, for a diagnosable reason, pointing straight at what would actually fix it. Measuring first is what made the second fix targeted instead of speculative.

  • FAISS
  • sentence-transformers
  • rank-bm25
  • cross-encoder/ms-marco-MiniLM-L-6-v2
  • SQLite
  • MCP