Build the core skills · 08 / 15

Make quality visible with an evaluation set

AI Engineer Path15 min guide

An evaluation set is a repeatable collection of tasks and expected behavior. It turns “this seems better” into evidence you can discuss. Write one before optimizing prompts or adding agents.

A manageable first evaluation

  1. Start with 20 development cases and 10 held-out cases. Grow toward 50–100 varied cases as the project matures. These are my practice targets, not industry standards.
  2. Include ordinary questions, missing evidence, ambiguous requests, malformed inputs and adversarial instructions embedded in documents. Label the expected behavior and supporting sources.
  3. Use code to check objective rules: valid schema, allowed category, existing source ID, prohibited tool not called.
  4. Use human review for factual support and usefulness. If you add a model grader, compare its decisions with your own labels and inspect disagreements.
  5. Record model, prompt, code and corpus versions. Run the same cases after a change and keep the failures, not just the average.
  6. Report held-out results once the approach is settled. If you repeatedly inspect and optimize against that set, it becomes development data; replace it with a fresh final set.

Your first scorecard

MeasureHow to report it
Task successPassed cases / total cases, with a written rubric
GroundingAnswers whose factual claims are supported by cited passages
AbstentionExpected refusals or clarifications handled correctly
RetrievalRecall@k on cases with labeled relevant documents
SpeedMedian and p95 end-to-end latency, with sample size
CostObserved usage and estimated cost per completed task
Critical failuresSeparate counts for unauthorized access or tool actions

Choose acceptance thresholds for your particular use case before comparing systems. A tiny hand-written set is a learning instrument, not proof of production reliability. Repeat unstable cases, report the denominator and uncertainty, and avoid using a single overall score to hide a serious failure.

For the underlying evaluation concepts, read Anthropic: Demystifying evals for AI agents ↗. The numbers, worksheet and exercises here are my suggested practice plan.

Move on when

A reviewer can rerun your evaluation, inspect the grading rules, see the baseline comparison and find at least three failures you can explain.

Your lesson resources

Download these files to follow along and put the lesson into practice.

Saved in this browser.