AgenticRAGFrameworks

AI Evals For Engineers & PMs

by Parlance Labs (Hamel Husain & Shreya Shankar) on Maven

IntermediateCoursePaid4 weeks, 3–5 hours per week (live cohort, recorded)

The reference cohort course on LLM evals: error analysis, trustworthy evaluators, and CI-integrated testing built around a real production agent.

Start LearningReviewed July 26, 2026

Overview

AI Evals For Engineers & PMs is taught by Hamel Husain, an ML engineer with 25+ years of experience, and Shreya Shankar, an ML systems and applied evals researcher. It runs as a live cohort-based course over 4 weeks at roughly 3–5 hours per week, with all sessions recorded for lifetime access; the next cohort listed is September 5 – October 3, 2026, and the list price is $4,200 USD (a 25% promotional discount was being offered at the time of writing). Coursework centers on building a production AI agent and then doing evals properly on it: instrumenting the agent for traceability, performing error analysis to find what is actually failing, designing evaluators you can trust, integrating testing into CI/CD pipelines, red-teaming for safety, and running cost-accuracy optimization experiments. The material spans evaluation strategies for different AI architectures — RAG systems, chatbots, multi-step agents and multi-modal systems — with techniques specific to each architecture's failure modes. Delivery includes 4 homework assignments, 10+ office hours, and a private Discord community. It is explicitly aimed at teams relying on manual spot-checking and at leaders who need visibility into failure patterns to prioritize work.

At a Glance

Topic
Agentic
Level
Intermediate
Format
Course
Cost
Paid
Duration
4 weeks, 3–5 hours per week (live cohort, recorded)
Provider
Parlance Labs (Hamel Husain & Shreya Shankar) on Maven
Hands-on
Yes — code/exercises
Certificate
None

What You’ll Learn

  • Instrument an AI agent for traceability so failures are diagnosable, not anecdotal
  • Run structured error analysis to identify and rank real failure modes
  • Design LLM evaluators you can actually trust, and validate that they agree with human judgment
  • Integrate evals into CI/CD so prompt and model changes are gated by tests
  • Red-team an AI product for safety issues
  • Run cost-accuracy optimization experiments to justify model and architecture choices
  • Adapt evaluation strategy per architecture — RAG, chatbot, multi-step agent, multi-modal

Highlights

  • Taught by Hamel Husain and Shreya Shankar, two of the most-cited practitioners on applied LLM evals
  • Built around shipping and measuring a real production agent, not toy examples
  • 4 homework assignments, 10+ office hours, private Discord, lifetime access to recordings
  • Material refreshed for the September 2026 cohort

Who It’s For

Best For

  • AI engineers shipping prompt or model changes without systematic measurement
  • Product managers who need visibility into AI failure patterns
  • Teams stuck on manual spot-checking of LLM outputs
  • Leads prioritizing where to spend effort on an AI product

Prerequisites

  • Experience building or shipping an LLM-powered feature
  • Working Python skills
  • Access to a real AI product or agent to apply the techniques to helps

FAQ

What is AI Evals For Engineers & PMs?

A live, cohort-based course from Hamel Husain and Shreya Shankar that teaches the discipline of evaluating AI products — error analysis, evaluator design, and continuous testing — by having you build and instrument a production AI agent. For engineers and PMs who ship prompt and model changes and currently have no systematic way to know whether they made things better.

Is AI Evals For Engineers & PMs free?

AI Evals For Engineers & PMs is a paid resource.

What level is AI Evals For Engineers & PMs for?

AI Evals For Engineers & PMs is aimed at a intermediate audience. Recommended background: Experience building or shipping an LLM-powered feature, Working Python skills, Access to a real AI product or agent to apply the techniques to helps.

How long does AI Evals For Engineers & PMs take?

Expect roughly 4 weeks, 3–5 hours per week (live cohort, recorded). Most learners work through it at their own pace.

What will I learn from AI Evals For Engineers & PMs?

You'll learn: Instrument an AI agent for traceability so failures are diagnosable, not anecdotal; Run structured error analysis to identify and rank real failure modes; Design LLM evaluators you can actually trust, and validate that they agree with human judgment; Integrate evals into CI/CD so prompt and model changes are gated by tests; Red-team an AI product for safety issues; Run cost-accuracy optimization experiments to justify model and architecture choices; Adapt evaluation strategy per architecture — RAG, chatbot, multi-step agent, multi-modal.

Topics

llm-evalserror-analysisagent-evaluationhamel-husainshreya-shankarcohort-course