Case study · RAG retrieval gate

How I'd stop a RAG regression reaching production

Retrieval changes are judged by eye more often than any other part of a RAG system. I built a gate that judges them on labelled questions, with statistics, before they merge.

The problem

Teams change the retrieval layer constantly: a new embedding model, a different chunk size, a hybrid retriever, a re-ranker. The failure that matters is quiet. The right passage slips from rank 2 to rank 12, drops out of the context window, and the model answers from something else. A spot check of five queries rarely catches it.

The opposite failure is just as costly. A gate with a fixed threshold fails good changes on noise, and after a few false alarms people stop trusting it and override it.

The gate design

Results

BEIR SciFact: 5,183 abstracts, 300 test queries, k=10, each candidate against BM25. Δ is candidate minus baseline.

Apple M1 Pro, CPU only, Python 3.13. Produced by scripts/results_table.py in the repository.
Configurationhit@10MRR@10 (Δ, 95% CI)nDCG@10 (Δ, 95% CI)Gatep50 / p95 ms
BM25 (baseline)0.81000.63210.6646baseline2.2 / 7.8
BM25, 12-word chunks0.73000.5316 (−0.1005, [−0.1354, −0.0661])0.5675 (−0.0971, [−0.1283, −0.0671])Fail, five metrics6.5 / 20.7
MiniLM-L6 dense0.79330.6047 (−0.0274, [−0.0689, +0.0151])0.6451 (−0.0195, [−0.0578, +0.0201])Pass, not significant8.9 / 13.7
Hybrid BM25 + MiniLM (RRF)0.82670.6536 (+0.0215, [−0.0041, +0.0466])0.6875 (+0.0229, [+0.0019, +0.0437])Pass, nDCG significantly better11.6 / 20.5
BM25 + cross-encoder re-rank0.81670.6527 (+0.0206, [−0.0136, +0.0562])0.6809 (+0.0163, [−0.0137, +0.0473])Pass1965.9 / 3923.8

The gate blocks the chunking change, which is worse well beyond noise, and passes the hybrid retriever, the only candidate with a significant gain. The re-ranker's gains are not significant on 300 queries and cost about two seconds a query on this CPU.

Trade-offs

Limits