Skip to content
Go back

What a RAG Eval Gate Actually Catches

I spent five years on the deterministic side of testing. An assertion either passes or it doesn’t. The element is on the page or it isn’t. The response equals the expected value or it fails. That’s the whole appeal: the test can’t be wrong about what it saw.

Then I built a small RAG app and tried to test it, and every one of those assumptions broke. The suite I ended up with is graded by another language model. Its results move run to run. It once told me an answer was missing a fact that was right there in the answer. And it still caught four real bugs I never would have found by hand.

This post is about those bugs. None of them were staged. All of them are the kind of thing a ten-question demo sails straight past and a real user hits three weeks later.

The setup, so the numbers make sense: a fictional product manual as the knowledge base, OpenAI embeddings and gpt-5-mini for retrieval and generation, DeepEval-TS for the evals (TypeScript-native, ships a Pytest-style runner that drops into CI), and GitHub Actions running the suite on every pull request. If a quality metric drops below its threshold, the check goes red and the merge is blocked. Full code: github.com/arminasbek/rag-eval-gate.

A RAG app can get worse while the tests stay green

A RAG app retrieves some documents, hands them to a language model, and asks it to answer from them. When it degrades, your CI doesn’t notice, because CI checks your code and this isn’t a code problem. Retrieval still runs. The model still responds. The unit tests still pass. The answers are just worse.

And “worse” is quiet. A prompt tweak, a model upgrade, a re-chunk of the corpus, a new batch of docs: any of them can silently drop answer quality that used to be fine. You find out when a customer complains. Or when they don’t complain, and just leave.

The gate turns that into a failing check. Same idea as a regression suite before a deploy, except it’s checking answer quality instead of button clicks.

The four checks

CheckWhat it asks
answer qualityIs the answer grounded in the retrieved context (faithfulness), and does it actually address the question (answer relevancy)?
retrieval qualityDid retrieval fetch the chunk the answer needs (recall), and rank it above the noise (precision)?
refusalFor a question the corpus can’t answer, does the app decline instead of guessing?
catches-hallucinationFed a deliberately wrong answer, does the gate reject it? A test of the gate itself.

Most of these are graded by another language model, an “LLM-as-judge.” Keep that in mind, it comes back at the end.

Watching it block a bad change

Before trusting the gate to catch subtle regressions, I checked it catches an obvious one. On a branch, I swapped the strict grounded-answer prompt for a generic “you are a helpful assistant.” Lint, typecheck, and the unit tests still passed. The eval suite didn’t:

The evals job failing in CI after the system prompt was weakened

Out-of-corpus questions started getting answered from general knowledge instead of declined, and the merge check went red. That’s the whole point of the thing, made concrete: a prompt change that looks harmless to every other check gets stopped here.

Now the real bugs, the ones I didn’t plant.

The app said “I don’t know” to questions it could answer

I promoted the suite from one hardcoded question to a set of 14 I knew were answerable, and ran it. Twelve passed. The two failures weren’t what you’d guess.

Both answers are in the corpus. The app refused questions it had the facts for.

This is worse than a hallucination, and quieter. When you picture a RAG failing, you picture a confident made-up answer, a wrong fact someone can spot and report. This is the opposite. There’s no wrong fact, no contradiction. The user asks, gets “I don’t know,” assumes the information doesn’t exist, and moves on. It reads as careful, safe behaviour, which is exactly why it survives a code review.

Watch how the two metrics split on it. Faithfulness scored both failures 1.00, a perfect score, because a refusal contradicts nothing in the context, so it’s technically faithful. Only answer relevancy caught them, because “I don’t know” isn’t a relevant response to an answerable question. If I’d only measured faithfulness, the axis everyone associates with hallucination, this bug would have passed clean.

So what broke? I looked at what retrieval actually returned for the weight question. The chunk with the weight in it never made the top four results. And when I read that chunk, the cause was obvious: at CHUNK_MAX_CHARS = 800, the chunker had packed three whole sections of the spec sheet into one blob.

## Storage and connectivity
The R7 has 256 GB of internal SSD storage. There is no removable SD card
slot. Data transfers over Wi-Fi 6, Bluetooth 5.2, or USB-C.

## Export formats
Scans export to the R7's native .rvz format, or to the .rvx interchange
format for third-party tools.

## Physical
The R7 weighs 1.15 kg with the battery installed and measures
180 x 95 x 62 mm. It carries an IP54 rating and operates in temperatures
from -10 C to 45 C. The current firmware version is 3.2.1.

Here’s the mechanism in plain terms. Each chunk gets turned into one embedding: a single list of numbers meant to stand for its whole meaning. This chunk is “about” storage, SD cards, Wi-Fi, Bluetooth, USB, .rvz, .rvx, weight, dimensions, IP rating, and firmware, all at once, so its embedding is the average of all of them. It points at nothing in particular. When the query is “how much does it weigh,” the one sentence that answers it is drowned out, and the chunk loses the ranking race to more focused ones. The fact is in the corpus. Retrieval just can’t find it.

The fix was smaller, single-topic chunks. I dropped CHUNK_MAX_CHARS from 800 to 200, so each ## section became its own chunk. Now the weight query retrieves just this:

## Physical
The R7 weighs 1.15 kg with the battery installed and measures
180 x 95 x 62 mm. It carries an IP54 rating and operates in temperatures
from -10 C to 45 C. The current firmware version is 3.2.1.

One topic, one focused embedding, ranks high for its own question.

The same thing showed up again later with “What warranty does it come with?” The warranty sentence was one line in a paragraph that also covered price, kit contents, and accessories:

The R7 base kit sells for 8,900 EUR and ships with the scanner, one
5,200 mAh battery, a USB-C charger, and a hard case. Corvid Instruments
covers the R7 with a 2-year limited warranty. A second battery and the
Field Charging Dock are sold separately.

Same fix, one level up: split the paragraph, give each fact its own heading, re-embed.

## Kit and pricing
The R7 base kit sells for 8,900 EUR and ships with the scanner, one
5,200 mAh battery, a USB-C charger, and a hard case.

## Warranty
Corvid Instruments covers the R7 with a 2-year limited warranty.

## Accessories
A second battery and the Field Charging Dock are sold separately.

The lesson: eval against a set of questions you know are answerable, and measure whether the app answered, not just whether what it said was grounded. A faithfulness-only gate lets silent refusals through, and no hallucination check will ever surface them.

Shrinking the chunks broke it a different way

Dropping the chunk size fixed the recall miss and introduced a new bug. Good warning about tuning one number without looking at what it produces.

At 200 characters, the splitter turned every document’s title line into its own chunk:

# Nyx R7 Technical Specifications
# Nyx R7 Troubleshooting
# Nyx R7 Portable LiDAR Scanner

Thirty-character chunks of pure noise. Except they contain the product name, and every question also contains the product name. So they score high similarity against basically any query and grab the top-ranked slots, pushing the chunk with the real answer out of the results. The weight question broke again, for a completely different reason than the first time.

The fix was structural, not another number: a heading should never be its own chunk. Fold a lone # Heading line into the block that follows it. Once headings ride along with their section body, the noise chunks disappear.

There was a cost to the 800-to-200 change I hadn’t thought about, too. More chunks means a bigger embedding index: the same text went from around 320 KB to around 870 KB, nearly tripled. Trivial here. Real money and latency at millions of chunks.

The lesson: a character count is a crude lever. It cuts on position, not meaning, so it happily produces garbage like a title-only chunk. For structured documents, split on the structure itself, the section headings, which is where I should have started.

While I was in here, the multi-topic chunks were also hurting precision: even when the right chunk was retrieved, it often ranked below less relevant ones. That pushed me to two-stage retrieval, a cheap first pass to gather candidates, then a small local reranker model to reorder them. Precision went from about 0.65 to about 0.90 with no change to recall. But that’s an improvement, not a bug the gate caught, so I’ll leave it there.

A correct answer scored red

After the chunking fixes, the suite failed a case that looks like the recall bug and is nothing like it. The question: “how much battery does the Nyx R7 need to run a firmware update?” The app answered:

The R7 must have more than 20% battery before a firmware update will start. Keep the scanner connected and do not power it off during the update.

That answer is correct. “More than 20%” is right there. Retrieval was fine, the chunk it pulled contained the requirement. Answer relevancy still scored it 0.33 and failed the 0.70 threshold.

Here’s why. Answer relevancy works by breaking the output into separate statements and scoring relevant-over-total. This answer has three: the battery requirement (answers the question), “keep the scanner connected” (doesn’t), “do not power it off” (doesn’t). One of three is on-question, so 1/3 = 0.33. The metric didn’t punish a wrong answer. It punished the two extra safety sentences, which came along in the same source chunk.

This red isn’t a bug in the app. It’s a real product tension the metric surfaced: relevance versus helpfulness. A terse answer scores high on answer relevancy. But for a product-support bot, “keep it connected, don’t power off during the update” is context a real user would want. Answer relevancy on its own is biased toward terse answers and against helpful ones.

And a second thing to notice. The judge’s written reason for the low score said the answer “omitted the specific battery requirement.” That’s false. The requirement is in the answer. The score had a mechanical basis, one of three statements, but the explanation the judge generated for it was simply wrong. Trust judge explanations at face value and they’ll occasionally lie to you.

The fix isn’t to lower the threshold until it passes, and it isn’t to cripple the bot into terse answers to satisfy one metric. It’s metric design. Answer relevancy is the wrong sole gate for “is this a good answer.” Pair it with a correctness check that verifies the output contains the right fact. Correctness passes this case (the fact is right) while still failing the silent refusals from the first bug (“I don’t know” isn’t correct). Relevance measures focus, correctness measures truth, and you want both.

The lesson: a red result is a question, not a verdict. Some reds are real bugs. Some are metric artifacts. Telling them apart, without blindly tuning thresholds until the dashboard is green, is the actual job.

The refusal test was flaky, like a UI test

The refusal check is the one place with no judge. It just compares strings, after normalising whitespace and case:

test.each(goldens("refusal"))(
  "refuses out-of-corpus question: $input",
  async ({ input, expectedOutput }) => {
    const { answer } = await askRag(input);
    expect(canon(answer)).toBe(canon(expectedOutput));
  },
);

expectedOutput for every refusal case is the same sentinel, “I don’t know based on the provided documents.” Deterministic on the assertion side. It still went red intermittently in CI.

The failure: “Does the Nyx R7 support 5G or cellular connectivity?” One run the app returned the refusal string and passed. Another run, same commit, it returned “No, the Nyx R7 does not support 5G or cellular connectivity.” That answer is fine, the corpus lists the connectivity options (Wi-Fi 6, Bluetooth 5.2, USB-C) and “no cellular” is a fair inference. It’s just not the exact string the test wants, so the test fails.

Two separate problems here.

The test case is borderline. A clean refusal case is one where the corpus says nothing at all: “Does it have GPS?”, “What colour options are there?”. This one is different, the corpus lists the connectivity, so a capable model can reason “the listed options don’t include cellular, therefore no.” That’s inference from the corpus, not a genuine gap. Borderline cases like this make the check flaky by construction. I’d already dropped an earlier one, “can the internal storage be expanded?”, for the same reason: the spec says “256 GB SSD, no SD slot,” which lets the model answer instead of refuse.

The other problem is deeper. Generated text is non-deterministic, so an exact-match assertion on it is inherently flaky. This is the same failure class as a browser test that asserts an exact on-screen string: it passes until the frontend rephrases a label. Here it passes until the model rephrases its refusal on a re-roll. Nothing pins determinism, and even at temperature zero these models aren’t perfectly repeatable.

I didn’t rewrite this one. The exact-string check is what’s in the repo, because for a small example it’s the most readable thing in the suite: no judge, no threshold, you can see exactly what it asserts. The cost is that it’s the least reliable of the four, and it’s flagged that way.

The lesson holds regardless: an LLM answer isn’t a deterministic output, so expect(sentence).toBe("...") against one is flaky by construction. The exact refusal string is the easy thing to assert, not the right thing. The property the check actually cares about is that the app declined and added no fact from the corpus. Testing that property, a contains-check plus a leak guard, or the LLM judge the other suites already use, would make it stable. It would also cost more than a string compare, which is the trade I left on the table.

The test tool is an AI too

This runs through all of it. In traditional test automation the assertion engine is completely deterministic. It can’t be wrong about whether an element is on the page. Here, the thing doing most of the grading is another language model, with its own failure modes:

The practical consequences: keep the CI gate small and strict, a few critical metrics, a hard threshold, blocks the merge. Keep any large exploratory eval run separate and non-blocking. A flaky judge should never block a merge on its own, which is also why you assert on aggregate pass rates over a set of questions rather than one pass/fail.

What this is worth to someone shipping RAG

A false refusal is invisible until a customer hits it. A hallucination at least leaves evidence. A refusal leaves nothing. The bot says “I don’t know,” the customer assumes the answer doesn’t exist, and they leave or quietly stop trusting it. It never lands in your logs as an error, because from the system’s point of view nothing went wrong.

You can’t eyeball your way out of this. A team demos the bot with ten questions, it answers all ten, ships. The failures live in the eleventh question, or in the one that breaks after next week’s prompt tweak. Manual spot-checking can’t cover that surface, and it definitely can’t cover it on every change.

What the gate sells is confidence, not code. It turns “we think the bot is fine” into “here’s the pass rate, and the merge is blocked if it drops.” The question that lands in a conversation about it: right now, if a change made your assistant start saying “I don’t know” to questions it used to answer, how would you find out? For almost everyone the honest answer is “a customer complains, eventually.” The gate changes that to “the pull request fails.”

What actually transferred

Coming from deterministic testing, I expected the assertions to carry over and they mostly didn’t. String equality is the wrong tool here. Exact-match is flaky. The grader is fallible. Half of what a traditional test suite relies on doesn’t hold.

What did transfer was the instinct underneath it: a green suite can be lying to you, and the job is to know how. The eval gate doesn’t remove that doubt. It just moves it somewhere you can see it, into a pull request check, instead of leaving it with the user who asked the eleventh question.

Code, corpus, eval suite, CI config: github.com/arminasbek/rag-eval-gate.


Share this post on:

Previous Post
Same Bot, Same Questions, Different Answer
Next Post
Deterministic Scripts Before Prompts: Cheaper, More Reliable AI Workflows