RAG, retrieval-augmented generation, is how you get a language model to answer from your own documents instead of its memory. I built a small one for my eval gate project: a fictional product manual, a chunker, two kinds of search, a reranker and a model. This post walks through each part in the order a question passes through it: what it does, why it’s there, and where it broke.
Some of the breaks were the obvious kind. The most surprising one wasn’t. For half the test questions, search ranked a chunk opening with a document title, like “Nyx R7 Technical Specifications”, above the chunk with the actual answer, because the questions and the title share the product name. The numbers in this post come from re-running the retrieval step by step and checking where the right chunk landed.

What RAG solves
Ask a general model about the Nyx R7 manual and it can’t know the answer. Its knowledge stops at a training cutoff, and a product manual it never saw isn’t in there at all. Worse, it usually won’t say so. It will produce a confident, plausible answer that is simply made up.
RAG fixes that by handing the model the right pages at question time. Retrieval finds the pieces of the manual most related to the question. Those pieces get pasted into the prompt as plain text, next to the question. The model answers from what it was just given, not from memory. Vectors only matter for the search step. The model itself never sees one.
Here’s what the model actually receives for the weight question, apart from its instructions: the four chunks retrieval picked, then the question.
Context:
[1] ## Physical
The R7 weighs 1.15 kg with the battery installed and measures 180 x 95 x 62 mm. It carries an IP54 rating (dust-protected and splash-resistant) and operates in temperatures from -10 C to 45 C. The current firmware version is 3.2.1.
[2] # Nyx R7 Technical Specifications
## Ranging and accuracy
The Nyx R7 has a measurement range of 0.3 m to 120 m. Range accuracy is plus or minus 3 mm at 50 m. ...
[3] ## Kit and pricing
The R7 base kit sells for 8,900 EUR and ships with the scanner, one 5,200 mAh battery, a USB-C charger, and a hard case.
[4] # Nyx R7 Portable LiDAR Scanner
The Nyx R7 is a handheld LiDAR scanner made by Corvid Instruments, first released in 2025. ...
Question: How much does the Nyx R7 weigh?
Only the first chunk answers it. The other three ride along because retrieval always keeps four, and two of them are the title chunks this post keeps coming back to.
Why not paste the whole manual?
Why not paste the whole manual into every prompt? For a real document set, three reasons. Size: a product’s docs, tickets and wiki run far past what fits in a context window. Cost: every token in the prompt is paid for on every question, including the 95% that has nothing to do with it. Quality: the more irrelevant text the model reads, the easier it is to miss the one sentence that matters, or to answer from a similar-looking wrong one.
To be straight about my own demo: the Nyx R7 manual is four short files, about 1,200 tokens. It would fit in a prompt easily. It stands in for a real doc set, so the pipeline behaves the way it would at scale.
Chunking
Retrieval doesn’t search documents, it searches chunks, so how you cut the manual decides what can be found. Too big, and one chunk covers several topics: its embedding mixes all of them and points at none, so a question about weight loses to more focused chunks. Too small, and you get fragments like a bare title line that match every question and crowd out the real answer. I hit both, in that order (full story). The fix was to split on the document’s own structure, its section headings, instead of a character count.

Embeddings and similarity
An embedding turns text into a vector, a long list of numbers that works like coordinates on a map. Texts with similar meaning land close together. To find the right chunk, the question gets embedded too, and cosine similarity compares its direction with every chunk’s direction: pointing the same way means similar meaning. The chunks are sorted by that score. In the simplest version of RAG, the top few go straight to the model.

Why vectors alone weren’t enough
That simplest version was my first one, and it was close but not right. Pure vector search found the right chunk for every question, inside the top four, but on 7 of 14 it ranked it below something else. This is where each question’s right chunk landed, first with vectors alone, then with each fix:
| Question | Vectors | BM25 | Hybrid + reranker |
|---|---|---|---|
| Operating temperature | 4 | 1 | 1 |
| Battery life | 3 | 2 | 1 |
| Weight | 4 | 17 | 1 |
| IP rating | 3 | 2 | 2 |
| Export formats | 2 | 1 | 1 |
| Warranty | 3 | 1 | 3 |
| Error code E03 | 3 | 1 | 1 |
| 7 other questions | 1 | 1 | 1 |
In all seven misses, the winner was the same kind of chunk: the opening chunk of a document, which carries the document title, like ”# Nyx R7 Technical Specifications” or ”# Nyx R7 Setup Guide”. Every question also says “the Nyx R7”. So those chunks matched every question on the product name, not on the answer. The chunk that actually said “The R7 weighs 1.15 kg” had no title and lost, 0.52 against 0.63. The fix from my last post, folding a lone title into the block after it, removed the tiny title-only chunks but not their pull. I haven’t proven this by re-embedding without the titles. It’s what the rankings point to.
Here’s the weight question with vectors alone, the top four and their scores:
| Rank | Score | Chunk |
|---|---|---|
| 1 | 0.626 | # Nyx R7 Technical Specifications / ## Ranging and accuracy |
| 2 | 0.585 | # Nyx R7 Setup Guide / ## First use |
| 3 | 0.553 | # Nyx R7 Portable LiDAR Scanner |
| 4 | 0.518 | ## Physical: “The R7 weighs 1.15 kg…” |
Three document openers, none about weight, all ahead of the one chunk that answers.
Hybrid retrieval and reranking
BM25 is an old keyword-search method, from long before embeddings, and it covers exactly where they’re weak. It scores chunks by which words of the question they contain, and rare words count for more. “R7” is in 19 of the 21 chunks, so matching it says almost nothing. “IP54” is in one, so a match on it counts heavily. So I run both searches side by side: the top eight chunks by meaning, the top eight by keywords, merged into one candidate list with duplicates removed, since a chunk can rank high in both. Neither list goes straight to the model. A reranker reorders the merged list first and keeps the best four.
The merge is a few lines in hybridRetrieve():
const seen = new Set<string>();
const candidates: Chunk[] = [];
for (const c of [
...retrieve(queryEmbedding, chunks, CANDIDATE_K),
...bm25Search(query, bm25, CANDIDATE_K),
]) {
if (seen.has(c.id)) continue;
seen.add(c.id);
candidates.push(c);
}
BM25 has its own hole, and the table shows it. It matches words exactly, with no stemming, so the question’s “weigh” never matched the chunk’s “weighs”, and the weight chunk fell to 17th.

The two searches are fast because neither one reads the question and a chunk together. Every chunk was embedded once, ahead of time. At question time only the question gets embedded and compared. That’s a bi-encoder: quick enough to scan everything, but each side is squeezed into a vector without seeing the other. A cross-encoder is the opposite. It reads the question and one chunk as a single input and outputs a relevance score, so it can see how the words of each relate. It’s more accurate, but nothing can be precomputed: every question and chunk pair is a full model run. So it only judges the shortlist. The fast searches gather up to sixteen candidates, the reranker (bge-reranker-base, running locally) scores them, and the top four go to the model.
It rescued the weight question from BM25’s 17th place. It isn’t perfect either: BM25 had the warranty chunk first, and the reranker moved it to third, still inside the four the model sees. With BM25 and the reranker added, mean Contextual Precision went from about 0.65 to about 0.90, with recall unchanged. I didn’t measure the two separately, so I can’t say how much each one contributed.
Generation
The model answers under a strict prompt: answer only from the given context, and if the answer isn’t there, reply with one exact sentence, “I don’t know based on the provided documents.” The guards in code return that same sentence when they block a question before the model, or catch an answer leaking the context after it. So every refusal looks identical, whoever made it, and a test can check for it with a plain string compare. That’s also its weakness: the model sometimes rephrases the sentence, and the check goes red. I covered that trade in the last post.
Evaluating retrieval
Two metrics grade the retrieval step itself, before the answer even matters. Contextual Recall asks whether the retrieved chunks contain everything needed for the correct answer. Contextual Precision asks whether the right chunks sit at the top, above the useless ones. Pure vector search passed recall on all 14 questions, because the right chunk always made the top four. It failed precision on 7, most likely because the right chunk was there but not first, sitting behind the title chunks. That’s the gap BM25 and the reranker closed together: recall stayed the same, mean precision went from about 0.65 to about 0.90. Both are graded by an LLM judge against an expected answer, which has its own problems. More on that in the last post.
The rank table above is a different, simpler measurement: no judge, just where the chunk with the answer landed. I re-ran it on the current version of the corpus while writing this post, so it shows the same pattern as the judged run, not necessarily the same seven questions.
Where this stops working
This pipeline is small on purpose, and each shortcut has a limit. The index is a JSON file, and every question is compared against every chunk. At 21 chunks that’s instant. At millions it isn’t, and that’s where a vector database comes in: an approximate index that trades a little accuracy for speed. The corpus is four short, made-up files with 14 test questions, so the 0.65 to 0.90 jump is a direction, not a benchmark. And questions go to search exactly as typed. With no stemming, “weigh” missed “weighs”. Real systems usually fix that by normalising words, or by having a model rewrite the question before searching.