Agent reliability
A section of the "Evaluating LLM applications and agents" topic. The questions check an understanding of practice, not memory of specific library APIs. After each answer comes a review: why the correct option is right, what is wrong with each incorrect one and which chapter to read.
Quiz difficulty Formats: one correct option, several correct options, order of steps, the mistake in an agent trace, pairs
01
What we check
- How Pass^k differs from Pass@k
- Which metric to report
- Analyzing the cause of a failure
- From a failure to a regression task
- Regression on a trajectory prefix
- Correct actions, wrong report
- Finding the error in a "silent" failure
- Rules first, then an LLM
- The agent's score dropped
- Judge the trajectory, not only the answer
- Protection against faked checks
02
Where to read
- Li, AI Agents in Depthsec. 7.2 · sec. 7.4.2 · sec. 7.5.2 · sec. 7.5.3 · sec. 7.6.1 · sec. 7.9
- Huyen, AI Engineeringch. 3 · ch. 6, Agent Failure Modes and Evaluation
- Agents and Vibe Codingch. 23