How to Evaluate Agent Output Quality

It seems to work is not an evaluation strategy. How to define quality concretely, build runnable evals, and track agent performance over time.

Published 2026-09-29 · 10 min read

It seems to work is not an evaluation strategy. Yet it is the strategy most agent deployments run on: someone eyeballs a few outputs, feels good, and ships. Then quality drifts — a model update, a prompt tweak, a new edge case — and nobody notices for weeks because nobody was measuring. Evaluating agent output is the unglamorous discipline that separates demos from dependable systems. This guide covers how to define quality concretely, build evaluations you can actually run, and track quality over time without turning your life into endless manual review.

Define done before you start

Every evaluation begins with acceptance criteria written before the agent runs, not after you see the output. What does a good result contain? What format must it follow? What must it never do? Write these as checkable statements: the summary names all three decision factors, not the summary is good. The discipline of writing criteria forces you to confront vagueness in the task itself — and vague tasks are the number one cause of bad agent output that was actually an unclear brief. Keep the criteria list short enough to check in under two minutes per output; if checking takes longer than doing the task, your evaluation will not survive contact with a busy schedule.

Build a small eval set

Collect ten to twenty representative tasks with known-good answers and run every agent change against them. Representative means covering the real distribution: the easy common cases, the tricky edge cases you have actually encountered, and at least one adversarial case designed to trip the agent up. Known-good answers do not need to be perfect — they need to be good enough that deviations are meaningful. Store the inputs, the expected outputs, and the date. When you change the prompt, the model, or the tools, rerun the set and diff the results. This is the cheapest regression protection available, and it catches the silent degradations — the model update that made outputs slightly worse across the board — that spot-checking misses.

What to score

  • Correctness. Does the output get the facts right? For verifiable tasks this is binary; for judgment tasks, compare against reference answers and note divergences.
  • Completeness. Did it do all of the task, or the easy eighty percent? Agents are skilled at producing confident partial work that looks finished.
  • Format compliance. Does it match the required structure, length, and conventions? Format failures break downstream automation even when the content is right.
  • Efficiency. How many steps and tool calls did it take? An agent that gets the right answer in forty steps is a cost and latency problem wearing a quality costume.
  • Safety and restraint. Did it stay within its permissions? Did it avoid irreversible actions, or quietly route around the approval gates?

Human review that scales

You cannot review everything, so review strategically. Sample outputs across the distribution: newest tasks, longest runs, tasks the agent flagged as uncertain, and a random slice of the routine ones. Review the sample against your acceptance criteria and log every failure with its category — wrong facts, incomplete work, format breakage, permission overreach. After a few weeks the failure log becomes the most valuable document you own: it tells you exactly where to tighten the prompt, add a check, or narrow the scope. Calibrate reviewers too: have two people score the same ten outputs occasionally, because good enough drifts between reviewers faster than you expect.

Automated checks catch the mechanical failures

  • Schema validation. If the output should be JSON, a table, or a fixed format, validate it mechanically. Format compliance is the easiest quality dimension to automate.
  • Factual spot-checks. For outputs with verifiable claims, script checks against your data sources — row counts, name matching, date ranges.
  • Diffing against references. For deterministic tasks, diff the output against the known-good answer and flag anything beyond cosmetic differences.
  • Model-as-judge, with caution. A second model scoring outputs against a rubric is useful for scale, but it inherits biases and misses the same things the generator misses. Use it for triage, never as the sole gate for high-stakes work.
  • Anomaly flags. Outputs far longer or shorter than usual, runs with unusual step counts, or tool-call patterns that deviate from the norm deserve human eyes.

Tracking quality over time

Quality is a time series, not a snapshot. Record eval scores and sample-review results with dates, and watch for trends: gradual drift after model updates, sudden drops after prompt changes, slow improvement as you tighten criteria. Keep a changelog of everything you change about the agent — prompts, models, tools, permissions — so you can correlate quality shifts with causes. When quality drops, the changelog tells you where to look; without it, you are debugging blind. Set a threshold for intervention in advance: decide now what failure rate triggers a rollback or a prompt revision, because deciding in the middle of an incident produces worse decisions.

Evaluation feels like bureaucracy until it catches the degradation nobody noticed. The teams running agents reliably all do some version of this: written criteria, a regression set, sampled human review, automated checks for the mechanical stuff, and trends over time. Start with the acceptance criteria and ten eval tasks — that afternoon of work pays for itself the first time it catches a silent regression before your users do.

Calibrating evals against human judgment

An eval rubric is only as good as its agreement with the humans whose judgment it replaces. Before trusting automated scores, calibrate: take twenty to thirty representative outputs, score them with your rubric (or your model judge), and independently have a knowledgeable human score the same set blind. Then compare. Where they disagree, the rubric is wrong, not the human — dig into each disagreement and ask what the rubric missed. Common gaps: the rubric rewards surface polish over correctness, it penalizes legitimate variation in approach, or its criteria are ambiguous enough that two humans would disagree with each other. Iterate until human-rubric agreement is high enough that you would act on the scores without re-checking. This calibration step is tedious and it is the difference between evals that measure quality and evals that measure something correlated with quality on good days.

  • Sample outputs across the quality spectrum — calibration on only good outputs teaches the rubric nothing about failure.
  • Score blind: the human judge should not see the rubric scores before judging, or anchoring corrupts the comparison.
  • Investigate every disagreement, not just the average agreement rate — the pattern of disagreements reveals rubric blind spots.
  • Re-calibrate when the task changes. A rubric tuned for summaries will misjudge code, and the drift is silent.
  • Document the calibration: what was tested, where it disagreed, what changed. Future you will need this when scores start looking odd.

Evals for non-deterministic outputs

Agents are non-deterministic: the same input produces different outputs across runs. A single sample tells you almost nothing about typical quality, and evals that score one output per case systematically mislead. The fix is sampling: run each eval case three to five times and score the distribution, not the instance. Track the mean quality, but also the worst case — for production use, the fifth-percentile output matters more than the average, because users remember failures, not averages. Pass-at-k metrics (did at least one of k attempts succeed?) suit tasks where you can retry in production; strict per-attempt metrics suit tasks where you cannot. And watch variance itself as a signal: rising variance across runs with no prompt change often means the underlying model was updated, and your evals just caught something your monitoring missed.

Diagnosing quality regressions

Scores dropped. Now what? Regressions have a small set of usual causes, and checking them in order saves hours. First, did the prompt or instructions change? Diff the current prompt against the last known-good version — including system prompts and tool descriptions, which people change without thinking of them as prompt changes. Second, did the model change? Providers update models silently; pin versions where possible and log which version produced each eval run. Third, did the eval set change? New cases, reworded cases, or a shifted sample can move scores without any change in the agent. Fourth, did the world change? Tasks grounded in current facts, APIs, or websites decay as reality moves on. Bisect ruthlessly: revert the most recent change and re-run the evals. The discipline here is treating evals as a scientific instrument — when the reading changes, suspect the instrument and the environment before concluding the agent got worse.

Budgeting your eval spend

Evals cost money and time — model-judge calls on hundreds of cases add up, and human review is the most expensive line item. Spend the budget where failure is costliest: full human review on high-stakes outputs (anything customer-facing, financial, or irreversible), sampled human review plus automated checks on routine work, automated checks alone on low-stakes drafts. A practical allocation for a team running agents in production: five to ten percent of your agent compute budget goes to evals, with the human-review portion concentrated on the newest capabilities and the most recent changes. And re-evaluate the evals themselves quarterly — an eval suite that never catches anything is either perfect or broken, and perfect is the less likely explanation. Evals are insurance: price them like insurance, not like a luxury.

Eval-driven development: tests before prompts

Borrow the red-green cycle from software engineering: write the eval before you write the prompt. Define five to ten concrete cases with expected outcomes, watch the current prompt fail them (red), then iterate on the prompt until they pass (green). This disciplines prompt work enormously — instead of tweaking wording by feel, every change is judged against the cases. It also prevents the most common prompt-engineering failure: improving the demo case while silently breaking three others, which the eval suite catches immediately. Keep the eval set small enough to run on every change and expressive enough to cover the task's real variety; a dozen well-chosen cases beat a hundred redundant ones. As the prompt matures, promote the trickiest production failures into the eval set — the suite grows into a regression shield. Teams that adopt eval-driven development stop arguing about prompt wording and start discussing case coverage, which is the conversation that actually improves quality.

  • Write cases with expected outcomes before touching the prompt — red first, then green.
  • Every prompt change runs the full suite: no silent regressions on cases you were not thinking about.
  • A dozen well-chosen cases beat a hundred redundant ones — optimize for variety, not volume.
  • Promote production failures into the eval set; the suite should grow from real mistakes.
  • Discuss case coverage, not wording preferences — the conversation that actually improves quality.

Keep reading