LLMs in production
What an application needs after launch: input and output guardrails, choosing and routing models, caching, cost and latency, structured output, degradation tests, monitoring and working with feedback. And an architecture-level decision: a prompt, RAG or fine-tuning.
Loading…
01The whole topic
Take the whole "LLMs in production" topic
Every question of the topic — one per fact, from easy to hard. An honest check of the whole topic rather than of a single section.
02
Topic sections
#QUIZQUESTIONSDIFFICULTYSTATUS
8.4Cost and latencyWhy an agent gets more expensive with every step, How to reduce the time to first token, How to reduce the total response time
8.5Self-checks and degradation testsHow to grow an architecture, Retries on output failures, Is an orchestrator needed
8.6Monitoring and driftQuality metrics for observability, Metrics designed around failures, What to log
8.7User feedbackImplicit signals in a dialog, Explicit and implicit feedback, Regeneration as a signal
8.8Prompt, RAG or fine-tuningRAG or fine-tuning for changing policies, RAG or fine-tuning, The order of adapting a model