Skip to main content
AgentsUse

Guide

Agent Harness Engineering: Where Tool Use Succeeds or Fails

The model gets the credit and the blame, but the harness decides what actually happens: which call runs, with what arguments, under whose permission, and what happens when it fails. A sourced guide to the layer, its six jobs, and the order to build them in.

Published2026-10-0814 min read

Take one tool (a web search API, say) and put it in front of two different agent setups. In the first, the agent calls it three times with overlapping queries, pastes 40,000 tokens of raw results into its own context, loses track of the original question, and answers a slightly different one. In the second, the same tool gets called once with a tight query, returns five results, the agent reads two pages and answers with citations. Same tool. Same model, often. The difference is the harness: the runtime layer that decides which call runs, checks its arguments, enforces its permissions, handles its failure, and decides what its result is worth.

Most tool-use failures blamed on models live in that layer. This guide walks the layer piece by piece: what it is, the six jobs it does, what the engineering teams who run agents in production have published about building it, and the order to build it in when you are starting from a bare loop of model-plus-tools.

The tool is rarely the problem

When an agent misuses a tool, the instinct is to swap the tool or upgrade the model. Sometimes that is right. More often the evidence points one layer down. Anthropic's engineering team, writing about the evaluations they run on their own agents, describes reading failure signals off the harness rather than the model: total tool calls, redundant calls, tool errors, and invalid-parameter errors. A spike in invalid-parameter errors, in their account, usually means the tool description is unclear. The model was not stupid. The interface was ambiguous, and the harness passed the ambiguity straight through to a real call.

Their earlier guidance on the agent-computer interface makes the same point from the design side: spend at least as much effort on how the agent uses tools as you would on a human interface. One of their examples has become the standard illustration. On their SWE-bench agent, the model kept making relative-path mistakes against a code-testing tool. They changed the tool to require absolute filepaths, a change that made the error structurally impossible, and the model used the tool flawlessly from then on. No model upgrade. No prompt pleading. A harness-level fix to the shape of the call.

This is why AgentsUse now runs a Harness pillar next to its directory. The directory answers which tools exist and how they are set up. The pillar answers what has to surround a tool before an agent can be trusted with it. The two questions fail independently: a verified, well-documented tool inside a careless harness still double-sends the email.

What the harness actually is

A harness is everything in the execution path that is not the model and not the tool: the dispatcher that turns a model proposal into a call, the validator that checks arguments against the tool schema, the policy engine that decides whether this call needs a human, the retry and timeout logic around the transport, the budgeter that trims results before they re-enter the context, and the log that records what happened. Frameworks ship machinery for all of it. LangGraph alone documents retry policies with defaults, approval interrupts with a decision vocabulary, durable checkpointing, and two kinds of timeout. Your harness is what you decide with that machinery.

A note on the word, because precision matters here. Anthropic's engineering posts, which are the best public documentation of this layer, do not use the word harness in the pieces this pillar draws on. They write about tool design and the agent-computer interface, about context engineering, about workflow patterns. We use harness as the umbrella term for one practical reason: a builder needs a name for the layer they own, the layer that is theirs to configure, test, and blame. Call it the runtime, the control layer, the glue. The jobs are the same.

There are six of them. The harness must constrain calls (check arguments before execution), budget results (decide what a tool's output is worth in context), gate the consequential ones (permissions and human approval), recover from failure (bounded retries, timeouts, cancellation, safe re-runs), record what happened (structured logs and grounded claims), and coordinate the sequence (order, parallelism, checkpoints). The patterns library on this site devotes a sourced page to each pattern under these jobs: the failure it prevents, how it works, when to soften it, and a checklist.

Design the interface like you would for a new hire

Anthropic's tools guidance reads, at first, like advice for tool authors: choose tools the model can use without contorting itself, namespace related functions, return meaningful names instead of opaque identifiers, keep responses token-efficient. Read it again and it is harness guidance, because the harness is where a team enforces these properties across tools it did not write. You cannot rewrite a vendor's API, but your harness chooses which of its endpoints the model sees, renames parameters at the boundary, and trims what comes back.

The token numbers deserve attention because they are the rare published ones. Anthropic describes a response_format choice in their examples where a detailed tool response ran 206 tokens and a concise one 72, roughly a third, with the detail (names, identifiers for downstream calls) reserved for the steps that need it. Their agent product caps tool responses at 25,000 tokens by default. Put together, the principle is that a tool result is spending from a fixed budget, and the harness, not the tool's default verbosity, should decide the amount. Context researchers on the same team named the underlying effect context rot: as the window fills, the model's grip on what is in it weakens. Every uncatalogued result you paste in spends some of that grip.

Format choices at the boundary pay compound interest. Absolute paths instead of relative ones. Enums instead of free text. Human-readable slugs instead of UUIDs, which Anthropic reports improved their retrieval precision because the model could reason about what it was holding. None of these require the vendor's cooperation. All of them belong to the layer you control.

What production runs add

Past interface design, the published production accounts converge on four additions. The first is checkpointing: save run state often enough that a failure resumes near where it broke. Anthropic describes pairing deterministic safeguards (retry logic, regular checkpoints) with model adaptability, and human-in-the-loop middleware in LangGraph resumes a paused graph from its checkpoint with the human's decision injected. The second is timeouts with real cancellation. The MCP specification, which governs a growing share of tool traffic, advises timeouts on all requests, defines a cancellation notification carrying a request id and a reason, and tells both sides to log when it fires. Latency without a ceiling is not patience; it is an unbounded liability.

The third is permission with vocabulary. A gate that only says yes or no teaches the model nothing. The richer decision set in current frameworks is approve, edit, reject, respond: run it, run this modified version, block it with a reason the model can learn from, or answer the agent with information instead of executing. The OpenAI Agents SDK adds a detail worth copying in any stack: when tool arguments are malformed, its approval flow fails closed, meaning the tool is not invoked and the call goes to a human. Unparseable goes to a person, never to the tool.

The fourth is evidence. MCP's host guidance reads like an audit checklist: show tool inputs to the user before calling, confirm before sensitive operations, validate results before passing them to the model, log tool usage for audit purposes. Anthropic runs full production tracing on its multi-agent research system and evaluates against a fixed set of example queries, watching decision patterns rather than reading conversation contents. The harness that produces that evidence on purpose, structured per call with outcome and cost attached, turns every incident from archaeology into arithmetic.

Six jobs, examined one at a time

Constrain comes first because it is cheapest. Every outgoing call gets checked against the tool's schema inside the harness, and rejections come back worded for the model. The MCP specification builds the whole pattern on a JSON Schema input definition per tool and puts validation on both parties; the harness is simply where a client honors its half. Teams skip this because the model usually gets the arguments right. The point is the tail: the one call in fifty that is malformed is the one that deletes the wrong directory.

Budget is the job teams discover latest and regret most. A tool result is not free text; it is a recurring charge against a finite window, re-read on every subsequent step. The mechanisms are unglamorous: a per-tool token cap (Claude Code's documented default is 25,000), response shapes the caller can size down (the concise-versus-detailed choice Anthropic measured at 72 against 206 tokens), pagination with steering text, and clearing results once the run has extracted their value. The failure this prevents has a name in Anthropic's context research, context rot, and it presents as forgetfulness, which is why it gets misdiagnosed as a model problem for months.

Gate is where consequence lives. The harness sorts tools into auto-run, ask-first, and never-autonomous, and the sort key is what a bad call costs, not how the tool is categorized in a store listing. Vendors already ship signals for this: MCP annotations mark tools read-only, destructive, and idempotent, and the GitHub MCP server can be launched read-only with a flag that overrides friendlier settings. A gate with only approve and deny starves the model; the frameworks converging on approve, edit, reject, and respond are solving a real teaching problem, not adding ceremony.

Recover covers everything between a failure and a decision. Bounded retries with error classification (LangGraph refuses to retry programming errors by default and retries server failures; that distinction is the whole pattern), timeouts with cancellation that both sides honor, and resume that restarts at the failed step rather than the first one. Anthropic's multi-agent system pairs this with checkpointing and telling the model the tool is failing, so it adapts instead of repeating. Recovery also includes the unsafe-replay problem: any step that restarts must know which of its effects already landed.

Record is the job that makes the other five debuggable. One structured line per call, args redacted, outcome classified, result size attached. From that single habit flow the metrics Anthropic publishes for its own agents (call volume, redundant calls, invalid-parameter rates per tool), the audit trail a permission gate needs to mean anything, and the evidence behind every claim a finished run makes. Grounding belongs to this job too: results validated before the model sees them, provenance kept attached, and uncited claims cut at write-up. A run you cannot replay on paper is a run you are operating on faith.

Coordinate decides what runs together and what waits. Dependencies between calls become data (this step consumes that step's output) instead of hopes embedded in a prompt. Independent calls fan out in parallel, which Anthropic credits with large time savings in its research system; dependent calls chain with gates between them, so a write never precedes the read that justifies it. This job is where harness design most resembles ordinary distributed-systems engineering, because it is ordinary distributed-systems engineering, with a language model as one of the unreliable components.

The build order

You do not build six jobs at once. Start where failures are most frequent and fixes are cheapest. Argument validation first: check every outgoing call against the tool schema in the harness, and return rejections worded for the model (which field, what shape, one example). It is an afternoon of work and it clears the most confusing failure class, malformed calls that execute anyway. Add a per-call timeout with cancellation at the same time; a call that cannot hang cannot stall a run overnight.

Next, classify your tools by consequence and gate the irreversible ones. Read the vendor's own signals first: MCP annotations mark tools read-only, destructive, or idempotent, and at least one major server sets them per tool. Gates go on sends, deletes, purchases, and permission changes, with the four-decision vocabulary, and malformed arguments route to a person by default. If the agent cannot yet be trusted with a tool ungated, the answer is a gate, not a better prompt.

Then the record: one structured log line per call, secrets redacted, outcome classified, result size recorded. This is what makes the rest tunable. Retry policy becomes tunable once you can see first-attempt failure rates per tool. Context budgets become settable once you can see which tools return 40,000 tokens a call. Sequencing and idempotency work lands last because it is the most design-heavy, and because the logs will show you exactly where duplicate effects and step races actually occur in your runs rather than where you imagined they might.

What a harness cannot fix

Three problems live outside this layer. A tool that returns wrong data quickly and consistently will defeat every pattern here except grounding, and grounding only catches shape errors, not falsehoods; that is a vendor problem, settled by switching tools, which is what a verified directory is for. An incoherent task stays incoherent under a perfect harness: if nobody can state what done looks like, gates and checklists merely formalize the wandering. And a tool with no safe repeat semantics and no idempotency support caps how autonomous the run around it can be. Naming those limits is part of harness engineering. The layer is powerful and it is not magic.

The thirty-day version

Compressed into a month, the build order looks like this. Week one: validation and timeouts. Wire schema checks into the dispatch path, set a per-call timeout from a week of observed latencies, and make cancellations real. The visible effect is immediate: malformed calls stop executing, hangs stop stalling runs, and both start appearing in the open as counted rejections and timeouts instead of mysteries. Week two: the consequence table and the gates. Sort every tool into auto-run, gated, and never-autonomous, attach the four decisions to the gated set, and make malformed arguments fail to a human. This is the week the agent stops being a liability with write access.

Week three: the record. One structured line per call, metrics aggregated per tool, and a standing monthly review with a name on it. Nothing about behavior changes this week, and everything about the following weeks does, because retry bounds, result budgets, and sequencing choices stop being guesses. Week four: spend what the data justifies. The tool with the worst invalid-parameter rate gets its boundary rewritten. The tool returning the fattest results gets a budget. The step that cannot be replayed safely gets its idempotency fix. A month in, the harness is not finished (it is never finished), but it has crossed the only line that matters: every common failure now has an owner, a mechanism, and a number.

What it can do, it does decisively. Most teams who instrument their runs find the same thing Anthropic found in theirs: the failures cluster in fixable places, unclear descriptions, unbounded results, missing gates, absent logs. Build the layer that watches those places, in the order above, and the tool you already have starts behaving like the tool you thought you bought.

Start smaller than this guide. Pick the one failure your runs produced this week (a malformed call that executed, a stall with no error, a result that ate the context), open the pattern page that owns it from the hub, and ship its checklist row. Harness engineering compounds because each fixed layer makes the next failure legible. The teams whose agents run unattended did not get there with a platform rewrite. They got there one owned failure at a time, with the receipts to prove each one stayed fixed.

One last framing, because it decides budgets. The model is rented; the harness is owned. Models will change under you this year, probably twice, and every harness property in this guide (the validation at the boundary, the gates, the budgets, the logs) survives the swap and usually works better on the stronger model. That is the economic argument for building the layer properly: it is the only part of the stack whose improvements you keep.

From the AgentsUse directory

Verified profiles for the tools and frameworks this guide discusses. Each passed the AgentsUse review gate: setup, limits, and maintenance signals checked against primary sources.

Continue in the Harness pillar

The pattern pages behind this guide: the failure each one prevents, how it works, and the checklist for shipping it.

Frequently asked questions

Is the harness just the framework by another name?

No. LangGraph, CrewAI, and the OpenAI Agents SDK ship harness machinery: retry policies, approval interrupts, checkpointing, guardrails. Your harness is the configuration and policy you build with that machinery: which tools run ungated, what the retry bound is, what gets logged, what a result may claim. Two teams on the same framework run different harnesses, and their agents behave differently because of it.

Which pattern should I build first?

Argument validation, then a per-call timeout. Both are an afternoon of work and together they remove the two most confusing failure classes: malformed calls that execute anyway, and calls that never come back. Add approval gates before the agent touches anything irreversible, and logs before it runs unattended. The production checklist guide orders all of it.

Can a better model substitute for harness work?

It changes the failure rate, not the failure list. A stronger model writes cleaner arguments and recovers from errors more gracefully, and Anthropic's own guidance is to spend effort on the tool interface first because capable models amplify good interfaces. A hung call still hangs, a duplicated side effect still duplicates, and an unlogged run is still unaccountable, whatever model sits inside.

Where do I check facts like retry defaults and token caps?

Every pattern page in the Harness pillar ends with a sources section and a checked date. The numbers used in this guide (LangGraph retry defaults, the 25,000-token tool-response cap in Claude Code, the concise-versus-detailed token comparison) come from the LangGraph documentation and Anthropic's engineering posts, fetched in October 2026.

Keep reading