Evaluating LLM applications and agents
How to tell whether an application got better after a change: evaluation datasets, LLMs as judges and their biases, agent reliability, RAG evaluation, statistical significance and observability in production.
Loading…
01The whole topic
Take the whole "Evaluating LLM applications and agents" topic
Every question of the topic — one per fact, from easy to hard. An honest check of the whole topic rather than of a single section.
02
Topic sections
#QUIZQUESTIONSDIFFICULTYSTATUS
6.1Evals and datasetsWhat to build an evaluation set from, labeling, linking metrics to business goals
6.5Statistics and productionSignificance of differences, observability, evaluation on live trafficSections: Statistical significance · Observability and tracing
6.6Benchmarks and their limitsWhat public benchmarks are for, Ranking by pairwise comparisons, An unexpected win in Elo