The Harness pillar
Good tools fail in bad harnesses.
A harness is the runtime layer between a model and its tools. It picks the call, checks the arguments, enforces the permissions, survives the failure, decides what the result is worth in context, and writes down what happened. This pillar collects the patterns for building that layer, sourced from the teams who run agents in production.
9 patterns · 3 new guides · 5 profiles with harness notes · sources checked 2026-10-08
The loop every agent run travels: the model proposes, the harness checks the proposal (arguments, permissions, budget), the tool acts, and the result comes back trimmed, logged, and grounded before it becomes context for the next step. Everything on this page hangs off that loop.
Between intent and execution
Six jobs, one layer
A note on the word. In the Anthropic engineering posts linked from this pillar, this layer is discussed without the word harness: they write about tool design, context engineering, and workflow patterns instead. We use harness as the umbrella term because a builder needs one name for the layer they own, and the layer does six jobs.
Constrain
Check every call against the schema before anything executes.
Recover
Retry what heals, stop on what cannot, resume where it broke.
Budget
Cap, shape, and clear tool results before they crowd out the task.
Gate
Let safe calls run. Hold the irreversible ones for a decision.
Record
Log every call and ground every claim in a result that arrived.
Coordinate
Run independent calls in parallel and dependent calls in order, with gates.
Why the layer matters
What the harness buys you
AgentsUse exists to improve how agents perform, stay secure, and stay reliable, not only to list their tools. Each outcome below names the failure it prevents, the patterns that do the work, and where AgentsUse tooling plugs in next. Every claim is a mechanism you can check, not a benchmark we invented.
Performance
3patterns
Prevents: Wasted steps: malformed calls failing late, bloated results, dependent calls guessed in parallel.
Calls are checked before they run, results are sized to the context they re-enter, and independent work fans out instead of queuing. Argument validation, Context budgeting, Sequencing.
RoadmapThe compatibility record puts the right tool in front of the model the first time, the cheapest performance gain there is.
Security
3patterns
Prevents: An action nobody approved executes, and nobody can reconstruct it later.
Only allowlisted tools run, destructive calls wait for approve, edit, or reject, and every call leaves an audit record. Permission gating, Tool-call logs, Grounding.
RoadmapThe permission and audit middleware ships these three patterns as a layer other builders can adopt.
Reliability
6patterns
Prevents: Runs stall on a hung call, loop on a flaky one, or repeat a side effect after a resume.
Retries are classified and capped, every call runs on a clock with real cancellation, and re-runs are made safe by design. Bounded retries, Timeouts and cancellation, Idempotency.
RoadmapReference patterns shipped as code, so recovery behavior is inherited rather than rediscovered team by team.
Cost and observability
2patterns
Prevents: Token burn nobody can attribute, and agent behavior nobody can review.
Result budgets cap what re-enters the context on every step, and structured logs make per-tool cost and failure rates readable at a glance. Context budgeting, Tool-call logs.
RoadmapThe audit layer turns the log into the monthly cost and behavior record for every connected tool.
One call, start to finish
Anatomy of a tool call
Zoom into a single pass through the loop. Validation and the permission check sit before execution on purpose: a call stopped there costs nothing. The failure branch is designed, not discovered: errors drop to a bounded retry, and past the bound they escalate to a fallback or a person instead of looping.
And every step leaves a record
{
"ts": "2026-10-08T14:03:11Z",
"run_id": "run_8f3",
"step": 7,
"tool": "web_search",
"gate": "auto-approved",
"args_hash": "b3:9f2c",
"duration_ms": 1840,
"outcome": "ok",
"result_tokens": 3120,
"attempt": 1
}Arguments are hashed, not stored. Outcome, duration, and result size are what the monthly review reads. The pattern: Tool-call logs and observability.
Start from the failure
The failure you keep hitting, and the pattern that ends it
Nobody browses patterns for fun. Find the symptom your runs actually produce, then read the one pattern that addresses it.
Without harness work
With the harness patterns
Same model, same tools. The left wire is what most demos ship; the right wire is the same calls with an owner for each failure mode. The table below maps each symptom to its pattern.
| Symptom in your runs | Pattern | Job |
|---|---|---|
| A flaky call loops forever, or a hopeless call gets retried until the budget is gone | Error recovery and bounded retries | Recover |
| Malformed arguments execute, or the model burns turns guessing the format you wanted | Argument validation before execution | Constrain |
| One verbose result evicts the instructions the model still needed | Context budgeting for tool results | Budget |
| The agent sends, deletes, or spends before anyone saw the plan | Permission gating and human approval | Gate |
| A retry or resume repeats a side effect: two charges, two emails, one edit applied twice | Idempotency and safe re-runs | Recover |
| Nobody can answer "what did the agent do, in what order, and why" after something goes wrong | Tool-call logs and observability | Record |
| One hung call stalls the whole run, and a cancelled tool keeps writing after the user said stop | Timeouts and cancellation | Recover |
| Calls that depend on each other race in parallel, or everything runs serially and the task takes ten times longer than it should | Sequencing dependent tool calls | Coordinate |
| The final answer asserts numbers and quotes that no tool ever returned | Grounding and citing tool results | Record |
The patterns library
Nine patterns, each sourced
Recover
Error recovery and bounded retries
A flaky call loops forever, or a hopeless call gets retried until the budget is gone.
Read the patternConstrain
Argument validation before execution
Malformed arguments execute, or the model burns turns guessing the format you wanted.
Read the patternBudget
Context budgeting for tool results
One verbose result evicts the instructions the model still needed.
Read the patternGate
Permission gating and human approval
The agent sends, deletes, or spends before anyone saw the plan.
Read the patternRecover
Idempotency and safe re-runs
A retry or resume repeats a side effect: two charges, two emails, one edit applied twice.
Read the patternRecord
Tool-call logs and observability
Nobody can answer "what did the agent do, in what order, and why" after something goes wrong.
Read the patternRecover
Timeouts and cancellation
One hung call stalls the whole run, and a cancelled tool keeps writing after the user said stop.
Read the patternCoordinate
Sequencing dependent tool calls
Calls that depend on each other race in parallel, or everything runs serially and the task takes ten times longer than it should.
Read the patternRecord
Grounding and citing tool results
The final answer asserts numbers and quotes that no tool ever returned.
Read the patternHow the frameworks score on these patterns. Twelve frameworks and agent runtimes compared cell by cell on permissions, approvals, recovery, logs, and tool-result context, sourced from their official docs. Paperclip and Goose included, with their roles named.
Compare frameworksThe AgentsUse framework is in the works. Roadmap, not a release: a harness layer built from the patterns on this page. Builders who want in early can ask.
Request early accessLong-form
The harness guides
Three new guides carry the pillar: the engineering overview, a failure-diagnosis walkthrough, and the production checklist. They build on the existing field guides on error recovery, tool permissions, and pre-flight checklists, which now cross-link back here.
Guide · 2026-10-08
Agent Harness Engineering: Where Tool Use Succeeds or Fails
The model gets the credit and the blame, but the harness decides what actually happens: which call runs, with what arguments, under whose permission, and what happens when it fails. A sourced guide to the layer, its six jobs, and the order to build them in.
Read the guideGuide · 2026-10-08
Diagnosing Tool-Call Failures: Read the Run, Fix the Layer
An agent run failed. The transcript will not tell you why. A field method for separating argument errors, tool errors, timeouts, context bloat, permission blocks, and sequencing races, with the fix that belongs to each layer.
Read the guideGuide · 2026-10-08
The Production Harness Checklist: 24 Checks Before an Agent Touches Real Work
Twenty-four pass-or-fail checks across the call, the gate, the run, and go-live, drawn from Anthropic's engineering posts, the MCP specification, and LangGraph and OpenAI documentation. Print it, score it, and widen scope only where every box holds.
Read the guideInside the directory
Harness notes, where the vendor docs support them
Five profiles now carry a harness-notes section: repeat behavior, result volume, scope rules, and polling behavior, each taken from the vendor's own documentation and dated. Profiles without a verifiable fact stay empty on purpose.
Tool profile
Harness notes: result volume is the budget lever
max_results accepts 1 to 20, and Tavily's agent guidance sets 5 as the starting point. Every result re-enters the context on each following step, so a high ceiling taxes the whole run, not just one call.
Open the profileTool profile
Harness notes: poll crawls to a terminal status
Crawl responses page results at 10 MB per response and hand back a next URL. The next URL also appears while pages are still processing, so treat job status (completed, failed, cancelled) as the end signal, never the absence of a next link.
Open the profileMCP profile
Harness notes: scope the server before the model runs
The --read-only flag (or GITHUB_READ_ONLY=1) restricts the server to read tools, and read-only wins over --tools selections: a wrapper cannot widen what the flag closed. That is permission gating enforced by the server instead of the prompt.
Open the profileMCP profile
Harness notes: snapshots are the context cost
The server works from structured accessibility snapshots rather than screenshots, which keeps each observation text-shaped and referenceable (click by ref) instead of image-priced.
Open the profileMCP profile
Harness notes: scope is dynamic, repeat behavior is per tool
Allowed directories come from the launch arguments or from client Roots. When Roots are present they replace the configured list entirely, and they can change while the server runs. Scope checks belong in the harness on every call, not just at connect time. With no allowed directory from either source, the server fails at startup.
Open the profileMore notes as the facts land
A note appears only when a vendor document states the behavior. Browse the tools directory and the MCP directory for the verified setup, pricing, and limits record behind each profile.
Roadmap · UpcomingNot shipped, not dated, not priced
The AgentsUse framework
The directory maps the tool layer. This pillar teaches the runtime layer around it. The next step is AgentsUse shipping that layer as a framework of its own, built from the patterns above. Three directions, all pulled from gaps the catalog keeps hitting:
- Compatibility truth: which tool works with which framework and model, answered from verified records instead of model memory.
- Permission and audit middleware: allowlists, approval checkpoints for destructive calls, and one structured log line per tool call.
- Reference harness patterns as code: validation, bounded retries, context budgets, and safe re-runs, shipped as components instead of advice.
This section is a roadmap statement. Nothing here is a shipped capability, and when something does ship it will be listed in the directory under the same verification rules as every other profile in it.
Mention the framework in your message. No date is promised; the early-access list hears first when there is something to run.
Questions
What is an agent harness?
The code and policy around a model that turns its intent into tool calls: it picks the tool, checks the arguments, enforces permissions, handles errors and timeouts, decides how much of each result stays in context, and logs what happened. The model proposes; the harness disposes.
Is the harness the same thing as the framework?
They overlap. A framework (LangGraph, CrewAI, the OpenAI Agents SDK) ships harness machinery such as retry policies, approval interrupts, and checkpointing. Your harness is the set of choices you make with that machinery: which tools run ungated, what the retry bound is, what gets logged. Two teams on the same framework run different harnesses.
Do I need all nine patterns?
No. Argument validation and a per-call timeout pay for themselves on day one. Add approval gates before the agent touches anything irreversible, logs before it runs unattended, and the rest as the failures show up. The checklist guide orders all nine by what breaks first.
Where does the directory fit in?
The harness assumes tools worth calling. The directory profiles verify setup paths, auth, pricing, and limits for the tools and MCP servers the harness drives, and five profiles now carry harness notes: vendor-documented facts about repeat behavior, result volume, and scope that a harness design needs.