Skip to main content
AgentsUse

Guide

Diagnosing Tool-Call Failures: Read the Run, Fix the Layer

An agent run failed. The transcript will not tell you why. A field method for separating argument errors, tool errors, timeouts, context bloat, permission blocks, and sequencing races, with the fix that belongs to each layer.

Published2026-10-0814 min read

An agent run failed, or worse, finished with a wrong answer, and the investigation starts with the transcript. The transcript is the worst witness available. It shows what the model said it was doing. It does not show which calls actually ran, what the tools returned, how long anything took, which results got cut, or where the run spent its money. Diagnosis from the transcript alone produces the two classic misreads: blaming the model for an interface problem, and blaming the tool for a harness problem.

This guide is a field method. It assumes you have, or will add, one structured record per tool call: timestamp, tool, outcome, duration, result size, attempt number. From that record, almost every tool-call failure separates into one of six layers, and each layer has a different fix. Fixing the wrong layer is worse than slow. It actively hides the problem: a bigger retry budget over a bad description teaches the model to make the same malformed call three times instead of once.

The six layers a failure can live in

Arguments. The call was malformed: wrong type, missing field, a format the tool never accepted. Detect it by counting invalid-parameter rejections per tool. Tool. The call was well-formed and the tool failed or returned junk: a 503, an empty result dressed as success, a changed response shape. Detect it from tool-error outcomes on valid arguments. Transport. The call never came back: hangs, stalls, late answers landing after the run moved on. Detect it from duration distributions and cancellations.

Context. The call succeeded and the run drowned in what it returned, or an earlier result got pushed out before it was used. Detect it from result-token totals per run and per tool. Permission. The call was blocked, gated, or never attempted because the credentials or scope were wrong. Detect it from gate decisions and authorization errors. Sequencing. Each call was fine and the order was wrong: a dependent call ran early, a resume replayed a side effect, parallel calls returned five result sets into one window. Detect it from the run graph, not any single line.

Symptom to layer: the quick map

  • Same malformed call retried repeatedly: arguments layer. The rejection message is failing to teach, or the format invites the mistake.
  • Calls succeed, answer cites figures no result contained: context or grounding failure. The model filled a gap left by a truncated or dropped result.
  • Run stalls with no error and no progress: transport layer. A call is hanging inside a timeout that does not exist or never fires.
  • Side effect happened twice after a pause or crash: sequencing layer, with an idempotency assist. The resume replayed an effect.
  • Agent asks permission for trivia but sails through the dangerous call: permission layer. Classification is by tool name, not by consequence and arguments.
  • Everything works in short tests and degrades in long runs: context layer. Results accumulate, instructions slide, the model starts repeating itself.
  • Failures cluster on one vendor tool across different tasks: tool layer. Stop fixing the harness and evaluate the tool.

Two of these signals come straight from published production practice. Anthropic's tools guidance lists the metrics they watch in agent evaluations: total tool calls, redundant calls, errors, invalid-parameter errors, and explicitly ties the last one to unclear descriptions. Redundant calls, many tools invoked to cover what one right-sized call should return, point at pagination and result-shape settings on the tool side and missing budgets on yours. Learn to read those four numbers per tool and most diagnoses take minutes.

Worked example one: the retry storm

Say a run shows one tool called nine times in ninety seconds, eight failures, then success. The transcript reads as persistence. The log record separates it. If the eight failures are invalid-parameter errors with the same field wrong the same way, this is the arguments layer: the model guessed a format, got a vague rejection, and guessed again. The fix is at the boundary: validate in the harness before dispatch, return an error naming the field and one valid example, and tighten the format itself where you can (absolute paths, enums). Retrying harder is precisely wrong here, because the call cannot succeed until its shape changes.

If instead the eight failures are 503s with rising waits between them, the transport and vendor are talking and the harness listened. That is the recovery pattern working: classified as transient, bounded, backed off. The diagnosis flips to capacity: eight retries to one success says the tool is sick or the run is mistimed, and the fix is scheduling, a fallback tool, or a vendor conversation, not harness surgery. Same symptom on the surface. Opposite prescriptions. The outcome class in the log is what tells them apart, which is why outcome classes (protocol error, tool error, timeout, ok) must be distinct fields and not a boolean.

Worked example two: the run that forgot the question

A research run answers fluently and slightly off-brief: a constraint from the original request is gone. Pull the result-token column. Typically one call, a broad search or a full-page browser snapshot, returned tens of thousands of tokens a third of the way in, and everything after reads like a new task. That is context rot doing exactly what Anthropic's context-engineering write-up describes: as the window fills, the model's hold on earlier content weakens, and the loudest recent material wins. The fix is budgetary, not intellectual: cap results per tool, default search tools to small result counts with raw content off (Tavily's own agent guidance starts at five results), use depth-limited snapshots for browser steps, and clear or summarize results once the run has extracted what it needed.

Verify the fix by rerunning the same task and comparing two numbers: tokens returned by tools, and whether the constraint survives to the final answer. If the answer improves while calls stay flat, you found the layer. If it does not, suspect grounding instead: the result was there, the model cited around it. That failure wants provenance attached to results and a cite-or-cut rule at write-up time, which is a different page in this pillar.

Worked example three: the double send

A notification went out twice, hours apart, from an otherwise healthy setup. Reconstruct the run graph around the send. The usual shape: the run paused for approval, the process restarted or the graph resumed, and the step containing the send executed again because its completion had not been recorded durably, or because the side effect sat before the approval point inside a step that restarts from the top on resume. LangGraph's documentation is blunt about that semantics: a node that suspends restarts from its beginning, so effects placed before the pause point run twice. The fix is placement and proof: side effects after the gate, completion written before the next step starts, and a check-then-write or idempotency key on the send itself. Then test it on purpose: pause mid-run, kill the process, resume, and count sends. One. Anything else is a failed fix.

Worked example five: the polite empty result

Every call succeeded. Durations normal, outcome column green, and the answer still wrong, because one tool returned an empty payload inside a success wrapper (a search with zero hits serialized as a result, a scrape that captured a cookie wall and called it content). The model treated absence as evidence and wrote around the hole. This is the grounding layer's signature failure: transport fine, shape wrong. The diagnosis is to compare what the answer cites against what the results contained. If a paragraph's claims have no corresponding payload, the run did not fail loudly, it failed politely.

The fix sits between the tool and the context: validate the result's shape before the model sees it, and turn empty or malformed success into a tool error the model must react to (retry with a narrower query, try the alternate source, or report the gap). MCP's host duties say this nearly verbatim: validate tool results before passing them to the LLM. Where the tool offers a structured output schema, validate against it and stop trusting prose wrappers. Then add the write-up rule: claims cite results, and a claim with no result behind it gets cut or flagged. Polite failures do not survive provenance.

When several agents share the blame

Multi-agent runs add one diagnostic wrinkle: the failure's location and its symptoms live in different processes. A worker returns a confident, wrong summary because its own context rotted three calls in; the orchestrator, seeing a clean summary, builds the plan on it. Read these runs bottom-up. Check the worker's per-call record first (calls, result tokens, outcomes), then the handoff: what exactly crossed into the orchestrator's context, and what got dropped on the way. Anthropic's multi-agent write-up is useful here precisely because it treats subagent outputs as engineered artifacts (summaries distilled before return) rather than transcripts forwarded whole. If your workers return raw exploration, the diagnosis and the fix are the same: the handoff is a harness component, and it needs a shape.

The other multi-agent trap is attribution. The orchestrator retries a failed subtask by spawning a second worker, which repeats the first worker's side effects. From the log this looks like two healthy runs. Only the run graph (which steps completed, under which run id) reveals it. Diagnose with the graph open, fix with durable step completion and worker outputs checked before re-dispatch. A retried subtask should resume or compensate, never silently redo.

Fixing at the right layer

Each layer owns a class of fix, and cross-layer fixes are the expensive habit to break. Arguments: tighten formats, validate early, write rejections for the model. Tool: repair or replace; verify replacements against the recorded failing runs, and read the directory profile's stated limits before blaming your code for a documented ceiling. Transport: per-call timeouts sized from real p99s, cancellation that both sides honor, late results quarantined. Context: caps, shapes, clearing, compaction. Permission: consequence-based classification with argument-aware predicates, and fail-closed behavior on anything unparseable. Sequencing: dependencies written as data, gates between dependent steps, parallel only where nothing is consumed.

One discipline ties the layers together: every fix ships with the probe that would have caught the failure. The retry storm gets an alert on invalid-parameter rate per tool. The forgotten question gets a per-run tool-token total on the dashboard. The double send gets a kill-and-resume test in the harness suite. Anthropic's evaluation practice is the grown-up version of the same instinct: a small fixed set of realistic tasks, run after every change, with the transcripts read. Roughly twenty examples caught real regressions in their setup. Yours can be smaller. It cannot be zero.

Verify the fix by replay, not by mood. Keep the failing run's inputs (the task, the tool set, the state it started from) and rerun them against the fixed harness. The symptom should be gone and the probe should show the mechanism working: the malformed call rejected with a teaching error, the cancellation logged with its reason, the fat result trimmed with its escape hatch intact. If the replay cannot be run because the task was never captured, that is itself the last diagnosis of the day: your runs are not reproducible yet, and check 19 of the production checklist (a fixed evaluation set) just became the fix.

Know when the diagnosis leaves your building. If the failure reproduces against the vendor's own examples, with well-formed arguments and a healthy network, you are holding a tool bug or an undocumented limit. File it with the failing call attached, and route around it in the harness (a wrapper that pre-validates, a budget where none was documented) while you evaluate the alternative. The directory profiles exist for this moment: check whether the replacement's stated limits already exclude your failing case before you migrate to it and meet the same wall wearing a different logo.

Triage by blast radius while you work. A failure a user can see (the wrong send, the duplicated charge) outranks a silent one (the degraded answer), which outranks an expensive one (the run that cost ten times the norm). Silent failures deserve the second slot, not the last: they compound invisibly, and every week one runs undetected it trains somebody to trust output that has not earned it. Severity decides the order of fixes. The layer method decides where each fix goes, and confusing those two decisions is how teams end up polishing the logging while the double send ships again. The runbook entry closes the loop: three lines (layer, probe, fix) that the next on-call inherits instead of re-deriving under pressure.

Worked example four: the gate that cried wolf

A team reports the opposite problem: nothing dangerous ever runs, because a human approves forty calls a day and has started approving without reading. The log confirms it. Median approval time is under two seconds, and the gated set includes reads, harmless lookups, and one genuinely destructive tool sharing a queue. This is the permission layer miscalibrated, and the failure is organizational before it is technical: approval attention is a finite budget, and spending it on trivia bankrupts it for the call that matters.

Diagnosis is quick once you look at gate decisions rather than call outcomes. Count approvals per tool per week, and read what the approver actually saw. If the gate displays a tool name but not the arguments, the human was decoration. The fix has three parts: reclassify by consequence so reads and reversible drafts run free, show the arguments (MCP's host guidance says to display tool inputs before calling, and for sensitive operations to confirm first), and keep the destructive tool gated with its arguments large on the screen. Measure again in two weeks. Approval time on the dangerous tool should go up, not down. That is the metric working.

The mirror failure also exists: a call that should have been gated ran free because classification was set by tool rather than by arguments. Reading every file in a repository is one risk; reading the credentials file is another, and they arrive through the same tool. This is why the diagnosis checks the gate's predicate, not the tool list. If your framework cannot express 'pause when the path looks like this', the harness must, even if it means a plain conditional in front of dispatch. The day the agent reads the secrets file unprompted, the tool list will not be what anyone blames.

Reading a run cold: the ten-minute protocol

When a failed run lands and nobody knows anything yet, work this order before opening the transcript. Minute one: pull the per-call records for the run and read only the outcome column. A clean column of ok outcomes moves you to context or grounding; a scatter of tool errors moves you to the tool; a single very long duration moves you to transport. Minute three: totals. Calls made, tokens returned by tools, wall-clock time, against your baseline for similar tasks. A run at three times the calls has a story in the middle of it.

Minute five: find the first anomaly, not the worst one. Failures cascade, and the loudest error is usually downstream of the first quiet one: the truncated result that made the model improvise, the gated call rejected without a reason that made it retry blind. Minute seven: read the transcript, now, and only around that point, to see what the model believed. The gap between what it believed and what the records show is the diagnosis in one sentence. Minute ten: write the layer, the probe, and the fix into the runbook before fixing anything, because the next failure will rhyme with this one and the note is how you will know.

The postmortem habit

Close every diagnosis by writing three things into the runbook: the layer, the probe now watching it, and the fix. Over a quarter this becomes the maintenance map of your harness: which tools eat context, which descriptions confuse models, which steps cannot be safely replayed. It is also the evidence base for the upgrades that actually cost money, a tool swap, a tracing platform, since each proposal can point at counted failures instead of vibes. The harness you can measure is the harness you can improve. Start with the six layers and one structured log line per call, and the next failed run will tell you where it hurts.

From the AgentsUse directory

Verified profiles for the tools and frameworks this guide discusses. Each passed the AgentsUse review gate: setup, limits, and maintenance signals checked against primary sources.

Continue in the Harness pillar

The pattern pages behind this guide: the failure each one prevents, how it works, and the checklist for shipping it.

Frequently asked questions

The agent finished and the answer looked fine. How do I know anything failed?

Check three numbers against your baseline for similar runs: total tool calls, tokens returned by tools, and wall-clock time. A run that quietly took three times the calls usually recovered from something by thrashing. If you have per-call logs, look for retries, tool errors the model routed around, and results that were truncated. Silent recovery is a cost signal even when the answer survives it.

Invalid-parameter errors keep spiking on one tool. Model or description?

Description, until proven otherwise. Anthropic reads exactly this metric as an interface-quality signal in their agent evaluations. Print the last twenty rejected argument sets. If the model keeps making the same shape of mistake, tighten the format (enums, absolute paths, examples in the error message) instead of enlarging the retry budget.

When is the right fix to switch tools?

When the failure is in the tool's data or semantics rather than in your layer: wrong results returned confidently, undocumented rate limits, repeat behavior that makes safe automation impossible. Your harness can survive a flaky tool; it cannot correct a lying one. Verify the replacement against the same failing runs before you swap, using the directory profile's limits and setup notes as the checklist.

Do I need full tracing to use this guide?

No. One structured line per call (tool, outcome class, duration, result size, attempt number) covers most of the diagnosis here. Full end-to-end tracing, which Anthropic runs on its multi-agent system, earns its cost when several agents or long chains make the failure's location ambiguous.

Keep reading