AI Evals: Everything You Need to Know
by Hamel Husain and Shreya Shankar
Answers to 60+ practical questions on evaluating LLM apps, from error analysis and LLM judges to agent and RAG evals
Overview
AI Evals: Everything You Need to Know is a question-and-answer guide on hamel.dev, written by Hamel Husain and Shreya Shankar, who also teach the AI Evals for Engineers & PMs course on Maven and wrote the O'Reilly book Evals for AI Engineers. The FAQ first appeared in mid-2025 and was republished in expanded form on September 18, 2026 (last modified September 21, 2026), with a PDF version. It holds more than 60 questions grouped into sections: getting started and fundamentals, error analysis and data collection, evaluation design and methodology, human annotation and process, tools and infrastructure, production and deployment, and domain-specific applications. The approach starts with reading traces. The authors recommend manually reviewing 30 or more traces before automating anything, letting error analysis decide which evaluators to write, preferring binary pass/fail judgments over 1-5 scales, and checking an automated LLM judge against human reviewers before trusting it. Later answers cover synthetic data generation and where it is unreliable, sampling production traces, evals on sensitive data, custom annotation tools, prompt versioning, guardrails versus evaluators, and how to evaluate RAG systems, coding agents, multi-turn conversations, human handoffs and multi-step agentic workflows. There is no code to run; it is a methodology reference that works with any stack or vendor.
At a Glance
- Topic
- ML
- Level
- Intermediate
- Format
- Guide
- Cost
- Free
- Duration
- Self-paced reference, 60+ questions in 7 sections; a few hours to read end to end
- Provider
- Hamel Husain and Shreya Shankar
- Hands-on
- No
- Certificate
- None
What You’ll Learn
- ✓How to run error analysis on traces and turn the failures you find into specific evaluators
- ✓Why binary pass/fail evaluations beat 1-5 rating scales and how to combine several into one metric
- ✓How to check whether an automated LLM judge agrees with human reviewers before trusting it
- ✓When synthetic data helps and the situations where synthetic test data becomes unreliable
- ✓How evals differ between CI/CD gates and production monitoring, and how often to run each
- ✓How to evaluate RAG systems, coding agents, multi-turn conversations and multi-step agentic workflows
- ✓What a custom annotation interface for reviewing LLM outputs should include and when to build one
- ✓How to sample production traces efficiently and re-run error analysis when a gold dataset goes stale
Highlights
- •Written by the instructors of the AI Evals for Engineers & PMs course and authors of the O'Reilly book Evals for AI Engineers
- •Expanded and republished in September 2026, with a downloadable PDF version
- •The original 2025 version reached 189 points and 43 comments on Hacker News, where practitioners shared it with their teams and also argued over its build-your-own annotation tool advice
- •Vendor-neutral and tool-agnostic: the methods apply whether you use Braintrust, Langfuse, Phoenix or a spreadsheet
- •Covers agent-specific problems such as large traces, human handoffs and multi-step workflows, which most eval tutorials skip
Who It’s For
Best For
- ✓AI engineers shipping LLM features who have no systematic way to tell whether a change helped
- ✓Technical product managers who need to set up error analysis and annotation with their engineers
- ✓Teams building RAG pipelines or agents who want a repeatable evaluation process before buying an eval platform
- ✓Tech leads deciding what an internal eval platform should standardize across teams
Prerequisites
- •Experience building or operating at least one LLM-powered feature
- •Familiarity with prompts, traces and basic LLM concepts such as RAG and tool calls
- •No coding required to follow the material
FAQ
What is AI Evals: Everything You Need to Know?
A free, long-form FAQ by Hamel Husain and Shreya Shankar on how to evaluate LLM applications in practice. It is written for engineers, technical PMs and tech leads who ship LLM features. After reading it you can run error analysis on traces, write binary pass/fail evaluators, validate an LLM judge against human labels and decide what belongs in CI versus production monitoring.
Is AI Evals: Everything You Need to Know free?
AI Evals: Everything You Need to Know is free to access.
What level is AI Evals: Everything You Need to Know for?
AI Evals: Everything You Need to Know is aimed at a intermediate audience. Recommended background: Experience building or operating at least one LLM-powered feature, Familiarity with prompts, traces and basic LLM concepts such as RAG and tool calls, No coding required to follow the material.
How long does AI Evals: Everything You Need to Know take?
Expect roughly Self-paced reference, 60+ questions in 7 sections; a few hours to read end to end. Most learners work through it at their own pace.
What will I learn from AI Evals: Everything You Need to Know?
You'll learn: How to run error analysis on traces and turn the failures you find into specific evaluators; Why binary pass/fail evaluations beat 1-5 rating scales and how to combine several into one metric; How to check whether an automated LLM judge agrees with human reviewers before trusting it; When synthetic data helps and the situations where synthetic test data becomes unreliable; How evals differ between CI/CD gates and production monitoring, and how often to run each; How to evaluate RAG systems, coding agents, multi-turn conversations and multi-step agentic workflows; What a custom annotation interface for reviewing LLM outputs should include and when to build one; How to sample production traces efficiently and re-run error analysis when a gold dataset goes stale.
Topics
Sources
This page was written from 4 sources, 3 on domains other than hamel.dev.