Skip to main content
AgentsUse

Harness pattern · Recover

Error recovery and bounded retries

Retry the failures that can heal, stop on the ones that cannot, and when retries run out, route somewhere deliberate instead of dying mid-task.

The failure it prevents

In production: A flaky call loops forever, or a hopeless call gets retried until the budget is gone.

Tool calls fail in production for boring reasons: a rate limit, a 503 from a vendor API, a socket that closed early. Without a recovery policy the run dies on the first one, and the user pays for a task that stopped at step two of nine. The naive fix is worse. Retry everything, forever, and a permanent error (a bad API key, a deleted resource) becomes an infinite loop that burns tokens and hammers a vendor who already told you no.

The third failure is quieter. The run retries, eventually succeeds, and restarts the whole task from the top because nobody saved progress. Side effects from the first attempt (the file written, the message sent) happen twice.

How it works

Classify before you retry. LangGraph publishes the clearest version of this rule: its default retry policy retries most exceptions but refuses programming errors such as ValueError and TypeError, and retries HTTP failures only on 5xx status codes. A programming error does not heal between attempts. A 503 often does.

Bound the retry with numbers, not hope. LangGraph ships defaults worth copying as a starting point: 3 attempts total including the first, half a second before the first retry, doubling each time (backoff factor 2.0), capped at 128 seconds, with jitter so a fleet of agents does not retry in lockstep.

When the bound is reached, do something designed. LangGraph runs an error handler after retries are exhausted, and the handler can route the graph to a compensation node (a Saga-style undo) instead of aborting. Anthropic describes the same pairing from the other direction: deterministic safeguards such as retry logic and regular checkpoints, plus telling the model the tool is failing so it can adapt its approach.

Keep the two error channels separate. In MCP, a protocol error (unknown tool, invalid arguments) is a JSON-RPC error, while a tool execution failure returns inside a normal result marked isError: true. The second kind is visible to the model and belongs in its context, worded so it can act on it. The first kind is your harness talking to itself.

The call, diagrammed

Anatomy of one tool callA proposed call passes validation and a permission check, then executes. Errors go to a bounded retry or escalate to a fallback or a human. Successful results are shaped to a context budget before entering the model context.RequestValidationPermission checkExecutionResult shapingContextErrorbounded retryEscalatefallback or humanfailure branch: retry inside the bound,then escalate instead of loopingevery step writes one line to the audit log

Where this pattern sits: on the failure branch, between execution and escalation. Also drawn on the Harness hub.

When to use it

  • Any tool call that crosses a network boundary or a vendor API.
  • Long runs where restarting from zero costs more than the retry logic does.
  • Tools whose failures you have seen heal: rate limits, 5xx responses, cold starts, lock contention.
  • Runs with side effects that need a compensation path when the main path gives up.

When to skip or soften it

  • Tools that are not idempotent and have no safe retry semantics; fix that first (see Idempotency and safe re-runs).
  • Validation failures and programming errors; retrying them wastes budget and hides the bug.
  • Interactive runs where a human is watching and would rather decide than wait out a backoff ladder.

Tradeoffs

Retries trade latency for completion. Three attempts with backoff can stretch one call past two minutes in the worst case, so the bound belongs in the same conversation as your timeout policy. Retries also mask flakiness: a tool that fails 40 percent of first attempts and always recovers looks healthy in outcomes and terrible in your logs. Log first-attempt failure rates separately, or you will never fix the tool that needs fixing.

A bounded retry policy (LangGraph)

LangGraph attaches a retry policy per node. The values below are the documented defaults from the LangGraph fault-tolerance docs, which are a sane starting point for most vendor calls.

A bounded retry policy (LangGraph)
from langgraph.types import RetryPolicy

retry = RetryPolicy(
    max_attempts=3,        # first attempt plus two retries
    initial_interval=0.5,  # seconds before the first retry
    backoff_factor=2.0,    # each wait doubles
    max_interval=128.0,    # no single wait exceeds this
    jitter=True,           # desynchronize a fleet of agents
)

graph.add_node("call_tool", call_tool, retry_policy=retry)

Shape source: LangGraph fault tolerance documentation. Placeholders only; adapt names and numbers to your stack.

Implementation checklist

  • Every external tool call has a written retry policy, even if the policy is "do not retry".
  • Errors are classified as transient, permanent, or programming errors before any retry runs.
  • Attempts are capped, waits back off exponentially with jitter, and the whole ladder fits inside the timeout budget.
  • Retry exhaustion routes to a designed outcome: compensation, a fallback tool, a human, or a clean stop with state saved.
  • Progress is checkpointed so recovery resumes near the failure instead of restarting the task.
  • First-attempt failure rate is visible in logs separately from final outcomes.

Sources and freshness

Claims on this page checked against these sources on 2026-10-08. Code and config blocks are shapes to adapt, not benchmarks.