Building Agent Error Recovery

Agents fail; the question is whether failure is graceful or catastrophic. A taxonomy of agent failures and the recovery patterns — retries, fallbacks, circuit breakers — that contain them.

Published 2026-09-29 · 10 min read

Agents fail. Tools error, plans go sideways, external systems change, and models make mistakes — often several of these in the same run. The question is never whether your agent will fail; it is whether failure is graceful or catastrophic. An agent that retries sensibly, degrades to partial results, and asks for help when stuck is a reliable tool. An agent that loops forever, corrupts state, or confidently produces garbage from bad inputs is a liability. This guide covers the failure taxonomy, the recovery patterns that contain failures, and how to make every failure teach the system something.

The failure taxonomy

  • Tool errors. APIs time out, return malformed data, or reject valid requests. The most common failure class and the most recoverable — usually.
  • Bad plans. The agent's approach was wrong from the start: wrong tool for the job, wrong order of operations, or a fundamental misunderstanding of the task.
  • Context loss. The agent forgot a constraint or an earlier finding because it fell out of the context window, and is now working from incomplete premises.
  • Permission blocks. The agent needs access it does not have — sometimes legitimately, sometimes because it is attempting something it should not.
  • External changes. The website redesigned, the API versioned, the file moved. The world changed under a workflow that assumed it would not.
  • Model mistakes. Confident wrong answers, misread instructions, arithmetic errors — the irreducible background radiation of working with language models.

Retry with intelligence

Retrying is the first recovery tool and the most abused. Retry transient failures — timeouts, rate limits, momentary unavailability — with backoff, not immediately, and not forever. Set a retry budget: three attempts is a sensible default, with the budget shared across the run so one flaky step cannot consume everything. Critically, distinguish retrying the same action from retrying with a changed approach. If the first attempt failed because the parameters were wrong, repeating identical parameters is not recovery, it is a loop. Instruct the agent to diagnose before the second attempt: what specifically failed, and what will be different this time? And log every retry with its reason — retry patterns are diagnostic gold, showing you exactly which tools and steps are unreliable.

Graceful degradation

  • Partial results over total failure. If three of five subtasks succeeded, deliver the three with clear notes on the two that did not — not a failure message for the whole run.
  • Fallback to simpler methods. When the sophisticated approach fails, try the simple one: cached data instead of live lookup, a template instead of generation, a human-readable summary instead of the full analysis.
  • Ask for help early. An agent that reports being stuck after two failed approaches is recovering; one that burns fifty steps flailing is not. Set the escalation threshold explicitly.
  • Preserve work in progress. Write partial results to the state file before attempting risky steps, so a failure does not destroy what succeeded.
  • Fail loudly, not silently. A clear error with context beats a plausible-looking result built on failed foundations — silent corruption is the worst outcome.

Circuit breakers

Borrowed from distributed systems, circuit breakers stop cascading failures: when a component fails repeatedly, stop calling it for a while instead of hammering it. For agents, this means tracking failure rates per tool and per workflow step, and tripping the breaker — pausing the run, alerting a human, switching to a fallback — when failures exceed a threshold. Without breakers, one failing dependency turns into an expensive retry storm: the agent burns its entire budget re-attempting something that was never going to work. Implement breakers at two levels: per-tool (this API is down, stop calling it this run) and per-run (this run has failed too many steps, stop and report). The breaker is not pessimism; it is the mechanism that turns a bad hour into a brief pause.

Post-failure learning

  • Log every significant failure with context: what was attempted, what failed, what the error was, and what the agent did next.
  • Categorize monthly. The categories reveal systemic issues: if permission blocks dominate, the setup is wrong; if bad plans dominate, the prompting is wrong.
  • Fix the prompt or the tool, not just the instance. Each failure category should produce a concrete change — a clarified instruction, a better error message, a narrower tool.
  • Add surprising failures to the eval set. The test suite should grow from real incidents; a failure that surprised you once should never surprise you twice.
  • Review near-misses. The run that almost corrupted data but did not is more informative than the clean runs — it shows where the guardrails were thin.

Error recovery is what separates agents you can leave alone from agents you must babysit. Build the taxonomy into your thinking, retry intelligently with budgets, degrade gracefully to partial results, breaker the cascades, and make every failure pay tuition in the form of a concrete improvement. Reliability is not the absence of failure — it is the presence of recovery.

Designing for partial failure

The most common design error is binary thinking: the run succeeds or it fails. Real agent work fails partially — seven of ten subtasks complete, the eighth hits a broken API, and the run's design determines whether you keep the seven or lose everything. Design for partial failure from the start: persist completed subtask results as they finish, so a later failure does not invalidate earlier work; make subtasks idempotent where possible, so retries do not duplicate effects; and define the minimum viable result — what subset of the work still delivers value if the rest fails. A research agent that returns seven solid sections and flags three it could not complete has succeeded usefully; one that returns nothing because section eight failed has failed completely. The difference is architecture, not luck.

  • Persist each subtask's result as it completes — never hold everything in memory until the end.
  • Make subtasks idempotent so retries are safe: check-then-act, not act-and-hope.
  • Define the minimum viable result: which subset of the work still delivers value alone.
  • Report partial results honestly — seven sections plus three flagged gaps beats a silent total failure.
  • Design the resume path first: a failed run should restart from its persisted state, not from zero.

Human-in-the-loop recovery

Some failures should not be retried automatically — they need a human judgment call. Define the handoff triggers explicitly: the agent is about to take an irreversible action after a failure, the failure involves data the agent cannot verify, the same subtask has failed twice and a third retry is just burning money, or the error message suggests a situation outside the agent's training. The handoff itself must be well-designed: the human gets the goal, what was completed, what failed, the error details, and the agent's recommended next step — a proper incident briefing, not a raw stack trace. And crucially, the human's decision must feed back into the run: approve the retry, modify the approach, or cancel cleanly, with the agent resuming from the decision point. A recovery handoff that dumps context on a human and restarts from zero is not a feature; it is an admission that recovery was never designed.

Recovery budgets: knowing when to give up

Every recovery strategy needs a budget — a point where the system stops trying and declares failure. Without one, retries compound: the agent retries the failed API, the retry triggers a fallback, the fallback fails differently, and an hour later you have spent real money achieving nothing. Set explicit budgets: maximum retries per subtask, maximum total run time, maximum cost per run. When the budget is exhausted, the run fails cleanly with a full report of what was attempted — which is itself a valuable artifact for debugging. The budget should scale with the stakes: a background research task gets a generous budget and graceful degradation, a customer-facing action gets a tight budget and fast escalation. Giving up is a feature, not a failure — the alternative is unbounded spending on a task that was never going to succeed.

Learning from incidents systematically

Every failure is a free lesson about your system's weaknesses, but only if you collect it. Run a lightweight postmortem for significant failures: what was the agent trying to do, what failed, what did recovery do, and what would have prevented it or recovered faster. Look for patterns across incidents rather than treating each as unique — three different failures that all trace to the same flaky tool are one problem, not three. Feed the findings back into the system: new regression tests, tighter validation, adjusted retry policies, better handoff triggers. The organizations with the most reliable agents are not the ones with the fewest failures; they are the ones with the tightest learning loop between failure and fix. Reliability is a process of converting incidents into immunity, one postmortem at a time.

Idempotency: the foundation of safe retries

Every retry strategy rests on one property: idempotency — the guarantee that repeating an operation produces the same result as doing it once. Without it, retries are dangerous: a retried payment charges twice, a retried record creation duplicates data, a retried notification spams the recipient. Design tools for idempotency from the start: accept client-generated operation IDs so repeats are recognized and deduplicated, check current state before acting (create only if absent, update only if changed), and make the check-and-act sequence atomic where the backend allows it. Where true idempotency is impossible — sending an email cannot be un-sent — use the outbox pattern: record the intent first, execute second, and reconcile the two on recovery. Then teach the agent the retry contract explicitly: which tools are safe to retry blindly, which require a state check first, and which must never be retried without human approval. Retries without idempotency are just errors with extra steps.

  • Client-generated operation IDs let backends recognize and deduplicate repeated calls.
  • Check-then-act: verify current state before mutating, so repeats converge instead of duplicating.
  • Outbox pattern for irreversible actions: record intent first, execute second, reconcile on recovery.
  • Classify every tool by retry safety: blind-retry, check-first, or never-without-approval.
  • Test retries explicitly — the retry path is production code, not an afterthought.

Keep reading