RAG Evaluation: Find the Failure Before the Fix
Evaluate RAG retrieval and answers separately. Build a small evidence set, measure missed sources, and use failure categories to choose the next improvement.
RAG evaluation becomes useful when it tells you what to change. A single “answer quality” score cannot explain whether your assistant searched the wrong documents, lost an important exception, or ignored evidence it already had. Start by separating retrieval from generation, then connect both to the user's actual task.
Consider an internal support assistant answering: “Can a workspace owner export deleted records?” An answer saying “owners can export data” sounds plausible. It is still wrong if the policy excludes deleted records. Your evaluation needs to notice that missing qualification before users do.
Build an evidence set before choosing metrics
For an initial, illustrative evaluation, collect 40 real questions that represent the decisions your assistant supports. This is a starting workload, not a statistical claim about sufficient sample size. For each question, record the answerable facts, relevant document versions, necessary passages, and acceptable abstention behavior. Keep sensitive customer information out of the fixtures.
Write the expected answer as requirements rather than one exact sentence. For the export example, require the response to distinguish active records from deleted records and to cite the applicable policy version. Several phrasings can satisfy those requirements. Matching a reference string would penalize harmless variation while missing a confidently worded omission.
Include questions with no answer in the corpus, outdated terminology, ambiguous product names, and evidence distributed across documents. A benchmark composed entirely of neat questions copied from section headings will tell you little about support tickets.
The original RAG paper combines a generative model with retrieved external memory. That separation motivates a practical debugging boundary: inspect what entered the prompt before changing how the answer is generated.
Measure whether the necessary evidence arrived
For a question with three labeled relevant passages, retrieving two gives passage recall of 2/3. That simple ratio is helpful only if the labels mean something. Ten overlapping chunks from one source should not count as ten independent pieces of evidence.
Define the unit you need: document, passage, or required fact. Document recall is useful for source discovery. Passage recall is stricter. Fact coverage can reveal whether retrieval found the policy rule but missed its exception. Report the definition beside the number.
For the export question, suppose the initial top five chunks contain the general export page but exclude the deletion policy. Record that as a retrieval miss. Increasing model size cannot reliably recover evidence your pipeline never supplied.
Rank also matters. A relevant passage in position 40 helps little when only the first six candidates enter the prompt. Record candidate recall and delivered-context recall separately. This distinguishes a search problem from a reranking or packing problem.
BEIR evaluates retrieval across heterogeneous datasets and found meaningful differences between retrieval approaches. Treat public benchmark results as evidence about those evaluated tasks, not a replacement for judgments on your own documents.
Grade answers against claims, not confidence
Break an answer into factual claims and ask three separate questions. Is each claim supported by supplied evidence? Does the answer satisfy the user's request? Are the citations attached to passages that actually establish those claims?
A response can be grounded but incomplete. “Owners can export active records” may be supported, yet fail to answer whether deleted records qualify. Conversely, a model might produce the correct policy from prior knowledge while citing an unrelated page. That is not dependable retrieval behavior.
Ragas proposes automated evaluation dimensions for retrieval-augmented systems. Automated grading can speed up review, but your product's acceptance criteria still need explicit definitions. Use human-reviewed examples to check whether a grader catches your important errors, including missing exceptions and unjustified certainty.
An illustrative result record could look like this:
{
"case": "deleted-record-export",
"necessary_evidence_found": false,
"answer_complete": false,
"unsupported_claims": 1,
"failure_stage": "retrieval"
}
These fields are a suggested local format, not a library API.
Run controlled comparisons
Freeze the corpus snapshot, questions, prompt, and generation settings while changing one retrieval variable. Compare the same cases before and after. Preserve per-case results so an improved average cannot conceal a regression on permissions or policy exceptions.
Run generation more than once on borderline cases when variability affects the decision. If one candidate wins only on one lucky run, report that uncertainty. For larger evaluations, calculate uncertainty using a method appropriate to your sampling process; do not imply that a tiny convenience sample represents every future user.
Track latency and expense alongside quality. A change that improves evidence coverage but doubles response time may still be useful for an asynchronous report and unacceptable for interactive support. The acceptance rule should reflect the interface.
Turn the report into a repair queue
Use failure categories with concrete owners: ingestion missed a document; permissions filtered incorrectly; search missed evidence; reranking displaced it; packing truncated it; generation contradicted it; or the expected answer was wrong. Inspect a handful of examples in every bucket before prioritizing work.
Our discussion of agent control flow provides the architectural backdrop for treating these as separate stages. The DB-GPT overview is related reading for assistants that connect language interfaces to organizational data.
Ship an evaluation change with the same care as a product change. Version the rubric, preserve previous results, and explain why labels changed. Otherwise, an apparent quality gain might simply mean the test became easier.
The useful output is a defensible next action: repair the missing deletion-policy ingestion, rerun affected cases, and check that the answer now includes the exception. A dashboard score is supporting evidence for that decision.