Tagged “debugging”
-
The Chunks Your Retriever Never Returns
Count how much of your index has ever been returned. The unreachable share is a defect list your eval set and your aggregate recall cannot see.
-
Reproducing a Bad Answer Somebody Sent You
A screenshot is not a bug report. Three replay modes that turn one complaint into a reproducible case, and what a failure to reproduce tells you.
-
Testing That Retrieval Respects Permissions
Every quality metric can look perfect while retrieval leaks documents across identities. Test it as an invariant with identities in the eval set.
-
Monitoring RAG Quality in Production
Live traffic has no labels. Eight signals you can watch anyway, plus scheduled canaries that catch config breaks within minutes.
-
The Five Places a RAG Pipeline Breaks
A bad answer implicates one of five stages. A bisection procedure that finds the responsible one in about ten minutes.
-
Why Your RAG Demo Works and Production Does Not
The demo and production differ in five specific ways, all of which favour the demo. Each one is measurable before you ship.
-
A Failure Taxonomy for RAG Answers
Read fifty bad answers, code each one, count the codes. The starter codebook and the procedure that turns anecdotes into a work queue.
-
Detecting Regressions When the Model Updates
A hosted model changes under you and your quality moves with no deploy. Version pinning, scheduled canaries, and the judge that drifts too.
-
Regression Testing a RAG Pipeline
A pipeline with no regression suite gets worse in ways nobody notices. Fixture corpora, range assertions, and must-pass cases in CI.