AI Evals
What are AI evals?
AI evals (short for evaluations) are methods for measuring whether an AI product or workflow is performing well. They give teams confidence their AI does what they expect.
Evals are not unit tests. A unit test expects the right answer every time. LLMs can give different answers to the same input, so an eval measures how often the LLM gets it right. An eval is a metric that counts how often a specific error occurs in your LLM output.
Which errors should an eval count?
The errors come from error analysis. Before picking an eval type, look at what the LLM is doing. Run a range of inputs through your workflow and log every mistake. The ones you most want gone are your candidates for evals.
You also have to define what a right answer looks like. Don't outsource this. Off-the-shelf eval tools bake in their own definitions of correct. Correctness is context dependent. Defining it for your product is the product team's job.
What are the four types of AI evals?
Once you know which error to count, pick the type that counts it best:
- Golden dataset — Known inputs with ideal outputs. Compare actual output to expected output. Best for small inputs with one right answer.
- Code assertions — Deterministic code that checks the output, like searching a transcript for a quote. Fast and cheap, start here.
- LLM-as-a-Judge — A second LLM judges the first LLM's output. Use it only when the error requires judgment.
- Customer feedback — The customer tells you if the response was good enough, through a rating or a behavior like editing the output.
In all four eval strategies, human graders should check how well the evals themselves perform.
Learn more:
- AI Evals: A Hands-On Guide for Product Teams
- Building My First AI Product: 6 Lessons from My 90-Day Deep Dive
- How I Designed & Implemented Evals for Product Talk's Interview Coach
- Behind the Scenes: Building the Product Talk Interview Coach
- 4 New Evals and 16 Experiment Variants to Fix One Customer Complaint
Related terms:
Last Updated: September 16, 2026