Learn / RAG in 7 lessons / Evaluating RAG
Evaluating RAG
Golden sets, faithfulness vs relevance, and why 'it looks right in the demo' isn't evaluation.
Why “it worked when I tried it” isn’t evaluation
A demo shows you the questions you thought to type. It doesn’t show you:
- questions phrased in ways your users actually phrase them, which are usually messier and less precise than a developer’s test queries,
- edge cases like a question with no good answer in your data at all,
- regressions - did your last chunking-size change make something that used to work now fail?
You need a fixed, repeatable set of test cases you can re-run automatically. That’s a golden set.
Building a golden set
A golden set is a list of {question, expected_answer_summary, expected_source_ids} entries.
Build it from:
- Real user questions, sampled from logs or support tickets if you have them.
- Deliberate edge cases: a question with a near-miss distractor (a chunk that’s topically
close but factually wrong for this specific question), an ambiguous question, and - crucially
- a question that has no answer anywhere in your data, to check the model refuses correctly instead of guessing (lesson 5).
- Regression cases: any real bug you find gets added as a permanent test case, so it can never silently come back.
30-100 well-chosen cases beats thousands of shallow ones. Quality of the edge cases matters more than raw count.
Two different things to measure
RAG has two stages, and each one fails independently, so measure them separately:
| Metric | Question it answers | What breaks it |
|---|---|---|
| Retrieval relevance | Did the right chunks even get retrieved? | Bad chunking, weak embeddings, a query phrased very differently from the source text (lesson 4) |
| Faithfulness / groundedness | Does the generated answer actually match what the retrieved chunks say? | The model paraphrasing loosely, mixing retrieved content with its own prior knowledge, or answering confidently despite weak context (lesson 5) |
A system can score well on one and badly on the other - retrieval can be excellent while the model still hallucinates a detail the context never stated, or retrieval can fail outright while the model still produces something plausible-sounding and “faithful” to nothing real.
Scoring at scale: LLM-as-judge, with a check on it
Manually grading 50+ answers every time you change something doesn’t scale. A common approach is LLM-as-judge: prompt a separate model call with the question, the retrieved context, and the generated answer, and ask it to score faithfulness and relevance against a rubric.
This works reasonably well, but it inherits the judge model’s own blind spots and can be fooled by confident phrasing the same way a human skimming quickly can. Treat it as a scaling tool, not a replacement for judgment: spot-check a sample of the judge’s scores against your own read of the same cases periodically, especially after changing the judge’s prompt or the underlying judge model.
Key takeaways
- A golden set is a fixed list of real questions with known-correct answers (and ideally, known-correct source chunks) that you re-run every time retrieval or prompting changes.
- Faithfulness (does the answer match what the retrieved context actually says) and relevance (did retrieval find the right chunks at all) are different failure modes and need to be measured separately.
- A demo only shows you the questions you thought to ask - a golden set should include edge cases: no-answer questions, ambiguous questions, and questions with a near-miss distractor chunk.
- LLM-as-judge scoring is useful for scale but needs periodic spot-checking against human judgment - it inherits its own blind spots.
Quick check
3 questions - see how much stuck.