Testing Strategies for AI Agents

Agents are non-deterministic, but untestable they are not. How to test at every level — from tool units to production canaries — without pretending agents are deterministic.

Published 2026-09-29 · 10 min read

Agents are non-deterministic, which tempts teams into the belief that testing does not apply. That belief is expensive. Untested agents fail in production in ways that are embarrassing, costly, or both — and the failures are usually in the parts that were perfectly testable. The trick is testing differently: not asserting exact outputs, but bounding behavior, checking distributions, and verifying the deterministic scaffolding around the non-deterministic core. This guide covers testing at every level, from individual tools to production canaries.

Why agent testing is different

Traditional software testing asserts exact outcomes: given this input, expect that output. Agents violate this contract by design — the same input can produce different valid outputs across runs, and some variation is not just acceptable but desirable. Testing therefore shifts from asserting specifics to asserting properties: the output contains the required elements, follows the required format, stays within bounds, and avoids the forbidden. You test that the agent's tools behave deterministically, that its prompts produce acceptable outputs across many runs, and that the system degrades gracefully rather than catastrophically. It is a different discipline, but it is a discipline — vibes-based shipping is not the alternative, it is the absence of engineering.

Levels of testing

  • Tool unit tests. Every tool the agent can call should have conventional unit tests: given these arguments, it returns the expected result or the expected error. Tools are deterministic code — test them like it.
  • Prompt tests in isolation. Run prompts against fixed inputs and check the outputs against your acceptance criteria across multiple runs. This catches prompt regressions before they reach the agent loop.
  • Scripted scenario tests. Drive the full agent through scripted scenarios with mocked tools, verifying it takes the right actions in the right order for known situations.
  • Eval sets on representative tasks. Ten to twenty real tasks with known-good answers, run on every change. The regression backbone of the whole practice.
  • Production shadow and canary. Run the new version alongside the old on real traffic — shadow for observation, canary for a small live slice — before full rollout.

Testing the non-deterministic core

For the parts that genuinely vary, test distributions instead of instances. Run the same task twenty times and check the pass rate against your criteria: does it succeed at least nineteen times? Do the failures cluster in a recognizable category? Fix the random seed where the framework allows it to make failures reproducible during debugging, then unfix it for the real evaluation — a test that only passes on one seed proves nothing. Define acceptable variation explicitly: for a summarization task, the key facts must appear every time while the wording may vary; for a classification task, the label must be stable. And test the edges deliberately: adversarial inputs, malformed tool outputs, and permission denials are where agents reveal whether their robustness is real or assumed.

Regression testing for prompts

Prompts are code and deserve the same regression discipline. Every prompt change — and every model change, which affects prompts silently — should trigger a run of the eval set with diffed results. Keep prompts in version control with the same review process as code; a casual prompt tweak can degrade behavior as thoroughly as a bad deploy. When a regression appears, bisect: revert halves of the change until the culprit is isolated. Maintain a library of historically tricky cases — the inputs that broke things before tend to break things again. And document why each prompt says what it says; prompts accrete clauses over time, and nobody remembers which clause fixed which incident without notes.

Load and cost testing

  • Step-count distributions. How many steps does a typical task take, and what does the tail look like? A task that usually takes eight steps but occasionally takes eighty is a cost incident waiting to happen.
  • Cost per task type. Measure real spend per completed task, not per month — monthly totals hide which tasks are wasteful.
  • Concurrency behavior. What happens when ten agent runs share tools, rate limits, or state files? Test it before production discovers it.
  • Timeout and stuck-run handling. Verify that runaway runs actually terminate — the kill switch is part of the system and needs testing too.
  • Degradation under load. When external APIs slow down, does the agent wait gracefully, retry sensibly, or spiral into expensive retry loops?

What good coverage looks like

Good agent test coverage is not a percentage; it is a set of questions you can answer confidently. Can you deploy a prompt change knowing within an hour whether quality changed? Can you show that the agent handles its ten most common failure modes? Do you know the cost and latency distribution of your main workflows? Can a new team member run the test suite and understand what is covered? If the answers are yes, your testing is working regardless of what the coverage number says. Review the suite quarterly: retire tests for workflows that no longer exist, add tests for failures that surprised you, and keep the suite fast enough that people actually run it.

Testing agents will never give you the certainty that testing a compiler gives you, and chasing that certainty is a trap. What it gives you is something more practical: the ability to change things without fear, to catch regressions before users do, and to know — with evidence rather than hope — how your agents behave. Start with tool unit tests and a ten-task eval set. That is an afternoon's work, and it is the foundation everything else builds on.

Property-based testing for agents

Example-based tests check specific cases; property-based tests check invariants that must hold across all cases. For agents, properties are often more valuable than examples because the input space is enormous. Define properties like: the agent never calls a write tool without first calling the corresponding read tool; every response to a customer includes either an answer or an escalation, never silence; the agent's final summary always mentions the files it changed. Then generate randomized inputs — varied tasks, edge-case phrasings, adversarial instructions — and verify the properties hold. Property tests catch the failures example tests miss: not the specific wrong answer, but the whole class of wrong behavior. They are harder to write, because stating the invariant precisely is real design work, but a handful of strong properties guard against more regressions than dozens of brittle examples.

  • State invariants precisely: never writes before reading, always answers or escalates, always reports what changed.
  • Generate randomized inputs: varied phrasings, edge cases, and adversarial instructions.
  • Test the property, not the example — one invariant covers thousands of cases.
  • Start with safety properties (what must never happen) before quality properties (what should happen).
  • When a property fails, the counterexample is the debugging starting point — save it as a regression test.

Testing tool interactions, not just outputs

Most agent tests assert on the final answer; the interesting failures happen in the middle. Test the tool-call sequence: did the agent check the documentation before calling the unfamiliar API, did it verify the record exists before updating it, did it stop after three failed attempts instead of looping forever? Record the full trace of tool calls in tests and assert on its shape — the order of operations, the absence of forbidden calls, the presence of verification steps. Mock tools at the boundary: a fake database that returns controlled edge cases, a fake API that fails on the third call, a fake file system with permission errors. The agent's behavior under these controlled adversities reveals whether your error handling is real or decorative. Final answers can be right for the wrong reasons; traces show the reasons.

Chaos testing: breaking things on purpose

Borrow from chaos engineering: deliberately inject failures into the agent's environment and observe. Slow tool responses to test timeout handling. Corrupt tool outputs to test validation. Revoke a permission mid-run to test graceful degradation. Feed the agent contradictory instructions from different sources to test prioritization. The goal is not to watch the agent succeed — it is to watch how it fails, because production will supply these conditions eventually and the failure mode is what you are actually shipping. Run chaos tests against staging with realistic tasks, and grade on the failure taxonomy: did it fail safely (stop and report), fail loudly (alert with context), or fail silently (the unforgivable one)? Every silent failure you find in chaos testing is an incident you prevented. Schedule these quarterly; systems drift, and last quarter's graceful degradation is this quarter's silent corruption.

Keeping the test suite alive

Agent test suites rot faster than traditional ones because the system under test changes constantly — models update, prompts evolve, tools get replaced. A suite nobody maintains becomes a museum of past concerns that passes while production burns. Assign ownership: someone specific is responsible for the suite's health, with time allocated, not squeezed between other work. Prune ruthlessly: tests that have not failed in six months and do not guard a critical property are candidates for removal — they cost maintenance and signal nothing. Add tests for every production incident, without exception; the incident review is not complete until the regression test exists. And version your tests alongside your prompts, so a prompt change automatically runs the relevant suite. The suite is a living document of what you have learned about your agent's failure modes — treat it with the same care as the agent itself.

Shadow mode and production replay

The highest-fidelity testing runs the agent against real production traffic without real consequences. In shadow mode, production inputs are duplicated to the new agent version, which executes fully — tools and all, against sandboxed or read-only backends — while the old version serves the actual user. Engineers then compare: where did the new version's behavior diverge, and were the divergences improvements or regressions? Production replay extends this: record weeks of real traffic, then replay it against candidate prompt or model changes to measure impact on realistic distributions rather than curated eval sets. Both techniques catch what eval suites miss — the weird inputs, the edge cases, the distribution shifts that only production exhibits. They require investment in sandboxing and traffic capture, but for high-stakes agent deployments they are the closest thing to knowing the future. Promote to production only what survived the shadow.

  • Shadow mode: duplicate production inputs to the new version; compare behaviors without user impact.
  • Sandbox the shadow's tools — full execution against read-only or fake backends, never production writes.
  • Replay recorded traffic against candidate changes to measure impact on real distributions.
  • Investigate every divergence: improvement or regression, and why the eval suite missed it.
  • Promote only what survived the shadow — the closest thing to knowing the future.

Keep reading