Datasets and evals
A section of the "Evaluating LLM applications and agents" topic. The questions check an understanding of practice, not memory of specific library APIs. After each answer comes a review: why the correct option is right, what is wrong with each incorrect one and which chapter to read.
Quiz difficulty Formats: several correct options, order of steps, one correct option
01
What we check
- Three sources of an evaluation set
- How traces become tests
- Slices and Simpson's paradox
- How to check that a bug is fixed
- Protecting a set from leaking into training
- Difficulty levels in a set
- A task with a user simulator
- A precise success condition
- The evaluation guideline
- Linking evaluation metrics to the business
- A repeatable environment for evaluating an agent
- Where to start an evaluation set
02
Where to read
- Li, AI Agents in Depthsec. 7.1.1 · sec. 7.3.1 · sec. 7.4.2 · sec. 7.4.3 · sec. 7.4.4 · sec. 7.4.5 · sec. 7.4.6 · sec. 7.8
- Huyen, AI Engineeringch. 4
- Agents and Vibe Codingch. 23