Building Software

Engineering Fundamentals for the Agent Era

Contents Section 9, Directing Agents

Evals: Testing AI Behavior

Mistakes to catch in review

  1. A prompt change shipped because it looked better on three examples.

  2. A model-graded score that disagrees with human judgment and was never checked against it.

  3. An eval set so easy that every version scores near perfect, so it cannot detect a regression.

Measuring nondeterministic systems with datasets, graders and statistics, so a prompt or model change is judged on evidence.

Topics

Eval Datasets
Building test cases with expected outcomes from real usage and past failures.
Deterministic Checks Before Model Judges
Using exact checks wherever possible, and model graders only for what code cannot decide.
Calibrating Model Judges
Measuring a model grader's agreement with human labels before trusting its scores.
Variance and Pass Rates
Running each case several times and reporting pass rates along with their uncertainty.
Evals as Regression Tests
Running evals on every prompt, model or tool change, the way tests run on every code change.

You understand it when you can

  • Explain what a twenty-case eval for a model-backed feature needs, and which of its checks can be deterministic.
  • Compare two prompt versions and decide whether the difference is real given run-to-run variance.
  • Explain how to check a model grader against human labels, and what agreement would make you trust its scores.

Drill

An agent reports that its new summarization prompt scores 4.6 out of 5 from a model judge on ten cases it chose itself, against 4.2 for the old prompt. Find every reason this does not show the new prompt is better, and design an eval that could.

Start here

Watch

How to Automate AI Evals (Correctly)

Hamel Husain, 2026. 27-minute explainer.

Argues for reading and labelling real traces before writing any rubric, so the eval set is built from the failures that actually happen instead of cases the author picked.

Watch

Read

AI Engineering: Building Applications with Foundation Models

Chip Huyen, 2025.

Two chapters on evaluation cover exact-match and similarity metrics, AI-as-a-judge and its failure modes, and how to design an evaluation pipeline for an open-ended model feature.

Trustworthy Online Controlled Experiments: A Practical Guide to A/B Testing

Ron Kohavi, Diane Tang and Ya Xu, 2020.

The standard text on variance, statistical power and false positives when comparing two versions, which is exactly the discipline a prompt A/B comparison needs.

Primary sources