Harness pattern · Recover
Timeouts and cancellation
Every call runs on a clock with a hard ceiling, cancellation is a protocol event both sides honor, and late results are ignored by design.
The failure it prevents
In production: One hung call stalls the whole run, and a cancelled tool keeps writing after the user said stop.
A tool call that never returns does not fail. That is the problem. The run sits there, the user watches a spinner, the budget meter runs, and every step after the hung call waits forever. Worse is the half-cancelled run: the user stops the agent, the agent stops talking, and the tool finishes its work in the background anyway, writing results nobody asked for into a system of record.
Cancellation also has a race. The caller gives up and moves on; the tool's answer arrives late and gets applied to a conversation that has already changed its mind. State built on late results is wrong in a way nobody can see.
How it works
Put a clock on every call. The MCP specification advises implementations to set timeouts for all requests even though it sets no default, and to send a cancellation notification when a request times out. Start from your p99 latency plus headroom, and write the number down per tool.
Treat cancellation as a protocol, not a hope. MCP's notifications/cancelled carries the request id and a reason. On sending it, the caller stops waiting. On receiving it, the callee stops work and frees resources. Late responses for a cancelled request must be ignored, and both sides should log the cancellation and its reason.
Keep a hard ceiling behind progress. Long calls report progress, and MCP progress notifications may reset the waiting clock, but the specification is explicit that a maximum timeout should always be enforced. LangGraph draws the same line with two different clocks: a run timeout (hard wall clock) and an idle timeout that refreshes when progress arrives. A node timeout surfaces as an error carrying which kind of timeout fired, so recovery logic can tell "hung" from "slow".
Cancel the context, not just the request. Timeouts should propagate: when the run dies, in-flight calls die with it, and the log records which calls were outstanding.
The call, diagrammed
Where this pattern sits: on the failure branch, between execution and escalation. Also drawn on the Harness hub.
When to use it
- Every networked tool call, without exceptions; a call with no timeout is a run with no ceiling.
- User-facing agents where a person can press stop and expects stop to mean stop.
- Long-running tools (crawls, scans, renders) where progress flows and the idle clock needs to differ from the total clock.
- Paid tools, where a hung call keeps spending.
When to skip or soften it
- Local, fast, in-process tools where the call cannot outlive the process anyway.
- Batch jobs deliberately designed to run to completion overnight with no interactive stop; keep the hard ceiling, drop the cancellation UI.
Tradeoffs
Timeouts cut real work that was about to finish, and a cut crawl or scrape wastes everything it had done so far. Mitigate with progress-based idle timeouts and partial-result handling where the tool allows it, rather than raising the ceiling until the pattern means nothing. Cancellation correctness also costs effort on the tool side; a server that ignores cancellation notifications forces the harness to compensate by ignoring late results and quarantining their effects.
Cancellation notification (MCP)
When a request times out or a user stops a run, the harness sends this notification with the request id and a reason. The receiver stops and cleans up; the sender ignores any late response for that id.
{
"jsonrpc": "2.0",
"method": "notifications/cancelled",
"params": {
"requestId": 42,
"reason": "Exceeded the 30s per-call budget"
}
}Shape source: MCP specification 2025-06-18: cancellation utility. Placeholders only; adapt names and numbers to your stack.
Implementation checklist
- Every tool has a written per-call timeout and a run-level hard ceiling.
- Progress resets the idle clock only; the maximum timeout fires regardless.
- Cancellation notifications carry the request id and a human-readable reason.
- Late responses for cancelled requests are dropped and logged, never applied.
- In-flight calls are cancelled when their parent run ends, and the log lists them.
- Timeout errors state which clock fired (idle versus total) so the fix targets the right one.
Sources and freshness
- MCP specification 2025-06-18: cancellation
- LangGraph fault tolerance (run and idle timeouts)
- MCP specification 2025-06-18: progress
Claims on this page checked against these sources on 2026-10-08. Code and config blocks are shapes to adapt, not benchmarks.