Case study · RAG retrieval gate
How I'd stop a RAG regression reaching production
Retrieval changes are judged by eye more often than any other part of a RAG system. I built a gate that judges them on labelled questions, with statistics, before they merge.
The problem
Teams change the retrieval layer constantly: a new embedding model, a different chunk size, a hybrid retriever, a re-ranker. The failure that matters is quiet. The right passage slips from rank 2 to rank 12, drops out of the context window, and the model answers from something else. A spot check of five queries rarely catches it.
The opposite failure is just as costly. A gate with a fixed threshold fails good changes on noise, and after a few false alarms people stop trusting it and override it.
The gate design
- Same questions, two configs. The baseline config is pinned to production; the pull request edits only the candidate. Both are scored on hit@k, MRR, recall, nDCG and, where labelled, required context, hard negatives and citations.
- Two conditions to fail. A metric fails only when the drop exceeds a practical threshold (0.02 absolute) and a 95% paired bootstrap interval over questions (10,000 resamples, fixed seed) excludes zero. The first stops trivial changes blocking a merge; the second stops noise doing it.
- Cost beside quality. Every run records index build time, p50 and p95 latency per query, peak memory and model size. These are reported, not gated, because they depend on the runner.
- No tuning on the test set. Fusion and re-rank settings are literature defaults fixed before the first run.
Results
BEIR SciFact: 5,183 abstracts, 300 test queries, k=10, each candidate against BM25. Δ is candidate minus baseline.
| Configuration | hit@10 | MRR@10 (Δ, 95% CI) | nDCG@10 (Δ, 95% CI) | Gate | p50 / p95 ms |
|---|---|---|---|---|---|
| BM25 (baseline) | 0.8100 | 0.6321 | 0.6646 | baseline | 2.2 / 7.8 |
| BM25, 12-word chunks | 0.7300 | 0.5316 (−0.1005, [−0.1354, −0.0661]) | 0.5675 (−0.0971, [−0.1283, −0.0671]) | Fail, five metrics | 6.5 / 20.7 |
| MiniLM-L6 dense | 0.7933 | 0.6047 (−0.0274, [−0.0689, +0.0151]) | 0.6451 (−0.0195, [−0.0578, +0.0201]) | Pass, not significant | 8.9 / 13.7 |
| Hybrid BM25 + MiniLM (RRF) | 0.8267 | 0.6536 (+0.0215, [−0.0041, +0.0466]) | 0.6875 (+0.0229, [+0.0019, +0.0437]) | Pass, nDCG significantly better | 11.6 / 20.5 |
| BM25 + cross-encoder re-rank | 0.8167 | 0.6527 (+0.0206, [−0.0136, +0.0562]) | 0.6809 (+0.0163, [−0.0137, +0.0473]) | Pass | 1965.9 / 3923.8 |
The gate blocks the chunking change, which is worse well beyond noise, and passes the hybrid retriever, the only candidate with a significant gain. The re-ranker's gains are not significant on 300 queries and cost about two seconds a query on this CPU.
Trade-offs
- Significance costs power. On a six-question fixture a real one-question regression reads as noise, so the small fixture in CI runs threshold-only as a tripwire.
- Passing is not improving. The dense MiniLM model passes because its drop is within noise, yet every point estimate is negative. A threshold-only gate would have failed it on MRR. If a change must prove it is better, read the significance verdicts, not the pass.
- Deterministic over clever. No LLM judge and no generator in the loop: the gate is free to run on every pull request and gives the same answer twice.
Limits
- One public dataset in one domain. Results on your corpus need your questions.
- It checks that labelled evidence is retrieved, not that an answer is true.
- Latency is one laptop CPU, one query at a time, pure-Python BM25. It ranks configurations; it is not a serving benchmark.
- The interval assumes the labelled questions represent real traffic. If they don't, it is precise about the wrong thing.