RAG evaluation
A section of the "Evaluating LLM applications and agents" topic. The questions check an understanding of practice, not memory of specific library APIs. After each answer comes a review: why the correct option is right, what is wrong with each incorrect one and which chapter to read.
Quiz difficulty Formats: several correct options, order of steps, one correct option, a phrase unsupported by the context
01
What we check
- Context precision and recall
- Why recall is harder to compute
- How faithfulness is computed
- What the RAGAS metrics measure
- Classes in answer correctness
- How to read RAG evaluation results
- What to evaluate in a RAG pipeline
- What goes into a question set for RAG
- Evaluating by component
- Order-aware metrics
- An invention in a support bot's answer
- A conclusion that is not in the data
02
Where to read
- Huyen, AI Engineeringch. 4 · ch. 4, Factual consistency · ch. 6
- Bratanič, Hane, Essential GraphRAGch. 8