LLM as a judge
A section of the "Evaluating LLM applications and agents" topic. The questions check an understanding of practice, not memory of specific library APIs. After each answer comes a review: why the correct option is right, what is wrong with each incorrect one and which chapter to read.
Quiz difficulty Formats: one correct option, several correct options, the faulty line, comparing answers against a rubric, sorting into groups
01
What we check
- What is not an LLM judge bias
- What belongs in a judge prompt
- Which scoring system to give a judge
- A judge is a model plus a prompt
- How not to go broke on a judge
- Judge biases and the defenses against them
- Rubric principles for a judge
- An agent and a judge from the same family
- A judge built from an expert's explanations
- What a judge can be
- Code or a judge
- An answer format for a program
- A check in code or by a judge
02
Where to read
- Li, AI Agents in Depthsec. 7.5 · sec. 7.5.1 · sec. 7.5.1, 7.5.4 · sec. 7.5.4
- Huyen, AI Engineeringch. 3 · ch. 3, How to Use AI as a Judge · ch. 4 · ch. 4, Create scoring rubrics with examples · ch. 4, Instruction-following capability
- Agents and Vibe Codingch. 23