Cost and latency
A section of the "LLMs in production" topic. The questions check an understanding of practice, not memory of specific library APIs. After each answer comes a review: why the correct option is right, what is wrong with each incorrect one and which chapter to read.
Quiz difficulty Formats: several correct options, one correct option
01
What we check
- Why an agent gets more expensive with every step
- How to reduce the time to first token
- How to reduce the total response time
- Degradation testing
- A fallback model
- Retries and latency
02
Where to read
- Li, AI Agents in Depthsec. 7.6.3
- Huyen, AI Engineeringch. 10
- Lakshmanan, Hapke, Generative AI Design Patternsch. 8, Pattern 27
- Gullí, Agentic Design Patternsch. 16