Skip to main content
AgentsUse

Harness pattern · Recover

Timeouts and cancellation

Every call runs on a clock with a hard ceiling, cancellation is a protocol event both sides honor, and late results are ignored by design.

The failure it prevents

In production: One hung call stalls the whole run, and a cancelled tool keeps writing after the user said stop.

A tool call that never returns does not fail. That is the problem. The run sits there, the user watches a spinner, the budget meter runs, and every step after the hung call waits forever. Worse is the half-cancelled run: the user stops the agent, the agent stops talking, and the tool finishes its work in the background anyway, writing results nobody asked for into a system of record.

Cancellation also has a race. The caller gives up and moves on; the tool's answer arrives late and gets applied to a conversation that has already changed its mind. State built on late results is wrong in a way nobody can see.

How it works

Put a clock on every call. The MCP specification advises implementations to set timeouts for all requests even though it sets no default, and to send a cancellation notification when a request times out. Start from your p99 latency plus headroom, and write the number down per tool.

Treat cancellation as a protocol, not a hope. MCP's notifications/cancelled carries the request id and a reason. On sending it, the caller stops waiting. On receiving it, the callee stops work and frees resources. Late responses for a cancelled request must be ignored, and both sides should log the cancellation and its reason.

Keep a hard ceiling behind progress. Long calls report progress, and MCP progress notifications may reset the waiting clock, but the specification is explicit that a maximum timeout should always be enforced. LangGraph draws the same line with two different clocks: a run timeout (hard wall clock) and an idle timeout that refreshes when progress arrives. A node timeout surfaces as an error carrying which kind of timeout fired, so recovery logic can tell "hung" from "slow".

Cancel the context, not just the request. Timeouts should propagate: when the run dies, in-flight calls die with it, and the log records which calls were outstanding.

The call, diagrammed

Anatomy of one tool callA proposed call passes validation and a permission check, then executes. Errors go to a bounded retry or escalate to a fallback or a human. Successful results are shaped to a context budget before entering the model context.RequestValidationPermission checkExecutionResult shapingContextErrorbounded retryEscalatefallback or humanfailure branch: retry inside the bound,then escalate instead of loopingevery step writes one line to the audit log

Where this pattern sits: on the failure branch, between execution and escalation. Also drawn on the Harness hub.

When to use it

  • Every networked tool call, without exceptions; a call with no timeout is a run with no ceiling.
  • User-facing agents where a person can press stop and expects stop to mean stop.
  • Long-running tools (crawls, scans, renders) where progress flows and the idle clock needs to differ from the total clock.
  • Paid tools, where a hung call keeps spending.

When to skip or soften it

  • Local, fast, in-process tools where the call cannot outlive the process anyway.
  • Batch jobs deliberately designed to run to completion overnight with no interactive stop; keep the hard ceiling, drop the cancellation UI.

Tradeoffs

Timeouts cut real work that was about to finish, and a cut crawl or scrape wastes everything it had done so far. Mitigate with progress-based idle timeouts and partial-result handling where the tool allows it, rather than raising the ceiling until the pattern means nothing. Cancellation correctness also costs effort on the tool side; a server that ignores cancellation notifications forces the harness to compensate by ignoring late results and quarantining their effects.

Cancellation notification (MCP)

When a request times out or a user stops a run, the harness sends this notification with the request id and a reason. The receiver stops and cleans up; the sender ignores any late response for that id.

Cancellation notification (MCP)
{
  "jsonrpc": "2.0",
  "method": "notifications/cancelled",
  "params": {
    "requestId": 42,
    "reason": "Exceeded the 30s per-call budget"
  }
}

Shape source: MCP specification 2025-06-18: cancellation utility. Placeholders only; adapt names and numbers to your stack.

Implementation checklist

  • Every tool has a written per-call timeout and a run-level hard ceiling.
  • Progress resets the idle clock only; the maximum timeout fires regardless.
  • Cancellation notifications carry the request id and a human-readable reason.
  • Late responses for cancelled requests are dropped and logged, never applied.
  • In-flight calls are cancelled when their parent run ends, and the log lists them.
  • Timeout errors state which clock fired (idle versus total) so the fix targets the right one.

Sources and freshness

Claims on this page checked against these sources on 2026-10-08. Code and config blocks are shapes to adapt, not benchmarks.