Statistical significance
A section of the "Evaluating LLM applications and agents" topic. The questions check an understanding of practice, not memory of specific library APIs. After each answer comes a review: why the correct option is right, what is wrong with each incorrect one and which chapter to read.
Quiz difficulty Formats: one correct option, several correct options
Quiz 6.5 "Statistics and production" also covers the section "Observability and tracing".
Quiz 6.5Statistics and productionStart practice Practice has no timer and a review after every answer.
01
What we check
- 73% versus 70%
- A paired comparison
- How many times to run
- An expected gain of 2–3 points
- Many hypotheses at once
- Is the set big enough
02
Where to read
- Li, AI Agents in Depthsec. 7.7
- Huyen, AI Engineeringch. 4
- Agents and Vibe Codingch. 23