Skip to main content
AgentsUse

Harness pattern · Recover

Idempotency and safe re-runs

Know which calls are safe to run twice, make the rest safe or unrepeatable, and arrange the graph so recovery never replays a side effect blindly.

The failure it prevents

In production: A retry or resume repeats a side effect: two charges, two emails, one edit applied twice.

Recovery mechanisms assume they can try again. Checkpoint resume re-runs a node, a retry policy repeats a call, a human approves a stalled run and the runtime restarts it. If the call inside writes data, "try again" can mean "do it twice": the payment captured twice, the notification sent twice, the text edit applied to already-edited text.

The trap is that tools differ and rarely say so plainly. The filesystem MCP server's documented hints make the spread concrete: create_directory is idempotent (creating an existing directory is a no-op), move_file fails outright if the destination exists, and edit_file is flagged destructive and not idempotent, because applying the same edit twice corrupts the file. Same server, three different answers to "what happens if this runs twice?"

How it works

Classify every tool once: safe to repeat, safe with a key, or never repeat blindly. MCP annotations give you idempotentHint; vendor docs and one controlled double-run test settle the rest. Write the classification into the harness config next to the retry policy, because the two decisions are one decision.

Place side effects after decision points. LangGraph documents the sharp edge: nodes restart from the beginning on resume, so any side effect placed before an interrupt() runs again when the node reruns. Move the effect after the approval, or make it conditional on state recorded in the checkpoint.

Rely on durable writes. LangGraph checkpoints successful node results (pending writes) so a resumed run does not re-execute nodes that already finished. Your harness should offer the same guarantee: a step's completion is recorded before the next step starts.

Where the vendor supports idempotency keys, send one per intended operation, generated by the harness. Where it does not, check before writing: read the current state and skip if the effect already landed.

When to use it

  • Every write tool that sits behind a retry policy or a resumable graph.
  • Payments, messages, publishes, and record creation, where duplicates are user-visible.
  • Long-running jobs that resume after crashes, deploys, or approval waits.
  • Multi-step sequences where a late failure triggers a re-run of earlier steps.

When to skip or soften it

  • Pure reads; repetition costs context budget, not correctness (see Context budgeting).
  • Vendor tools with documented idempotent semantics, verified once and recorded.
  • Throwaway sandboxes where duplicate effects are deleted with the sandbox.

Tradeoffs

Making a call idempotent usually means an extra read before the write, or a key store to maintain, and it slows the happy path. The alternative is a cleanup story for duplicates, which is slower and sometimes impossible (you cannot un-send). Cost it plainly: idempotency work scales with how bad a duplicate is, and for money and messages that answer is "very", so pay it there first.

Risk hints on a real tool set (filesystem MCP server)

MCP annotations let a server declare how a tool behaves on repeat. The filesystem server's documented hints, fetched from its README, show why one retry rule cannot fit every tool.

Risk hints on a real tool set (filesystem MCP server)
{
  "create_directory": { "idempotentHint": true,  "destructiveHint": false },
  "edit_file":        { "idempotentHint": false, "destructiveHint": true  },
  "move_file":        { "idempotentHint": false, "destructiveHint": false },
  "read_file":        { "readOnlyHint": true,    "idempotentHint": true  }
}

Shape source: Filesystem MCP server README (tool annotations). Placeholders only; adapt names and numbers to your stack.

Implementation checklist

  • Every write tool has a recorded repeat-behavior class: idempotent, keyed, check-then-write, or never-repeat.
  • Retry policies and resume logic consult that classification before repeating anything.
  • Side effects sit after approval interrupts, never before them, inside a resumable node.
  • Step completion is durable before the next step starts, so resume skips finished work.
  • A deliberate double-run test was executed once per write tool and the result written down.

Sources and freshness

Claims on this page checked against these sources on 2026-10-08. Code and config blocks are shapes to adapt, not benchmarks.