Benchmarks and their limits
A section of the "Evaluating LLM applications and agents" topic. The questions check an understanding of practice, not memory of specific library APIs. After each answer comes a review: why the correct option is right, what is wrong with each incorrect one and which chapter to read.
Quiz difficulty Formats: several correct options, one correct option
Quiz 6.6Benchmarks and their limitsStart practice Practice has no timer and a review after every answer.
01
What we check
- What public benchmarks are for
- Ranking by pairwise comparisons
- An unexpected win in Elo
- A new model is better on benchmarks
- Latency metrics when choosing a model
- A short budget and long tasks
02
Where to read
- Li, AI Agents in Depthsec. 7.5.4 · sec. 7.6.1 · sec. 7.6.4
- Huyen, AI Engineeringch. 4