Skip to main content
AgentsUse

The Harness pillar

Good tools fail in bad harnesses.

A harness is the runtime layer between a model and its tools. It picks the call, checks the arguments, enforces the permissions, survives the failure, decides what the result is worth in context, and writes down what happened. This pillar collects the patterns for building that layer, sourced from the teams who run agents in production.

9 patterns · 3 new guides · 5 profiles with harness notes · sources checked 2026-10-08

The harness loopThe model proposes a tool call. The harness validates arguments, checks permissions, and budgets context before the tool acts. The shaped result returns as context. A human checkpoint can approve, edit, or reject consequential calls.Modelproposes the next actionHARNESSthe control layerValidate argumentsPermission gateContext budgetRetry + timeoutAudit logToolsAPIs, MCP servers, codeContextwhat the model sees nextshaped result, within budgetHuman checkpointapprove · edit · reject

The loop every agent run travels: the model proposes, the harness checks the proposal (arguments, permissions, budget), the tool acts, and the result comes back trimmed, logged, and grounded before it becomes context for the next step. Everything on this page hangs off that loop.

Between intent and execution

Six jobs, one layer

A note on the word. In the Anthropic engineering posts linked from this pillar, this layer is discussed without the word harness: they write about tool design, context engineering, and workflow patterns instead. We use harness as the umbrella term because a builder needs one name for the layer they own, and the layer does six jobs.

Constrain

Check every call against the schema before anything executes.

Recover

Retry what heals, stop on what cannot, resume where it broke.

Budget

Cap, shape, and clear tool results before they crowd out the task.

Gate

Let safe calls run. Hold the irreversible ones for a decision.

Record

Log every call and ground every claim in a result that arrived.

Coordinate

Run independent calls in parallel and dependent calls in order, with gates.

Why the layer matters

What the harness buys you

AgentsUse exists to improve how agents perform, stay secure, and stay reliable, not only to list their tools. Each outcome below names the failure it prevents, the patterns that do the work, and where AgentsUse tooling plugs in next. Every claim is a mechanism you can check, not a benchmark we invented.

Performance

3patterns

Prevents: Wasted steps: malformed calls failing late, bloated results, dependent calls guessed in parallel.

Calls are checked before they run, results are sized to the context they re-enter, and independent work fans out instead of queuing. Argument validation, Context budgeting, Sequencing.

RoadmapThe compatibility record puts the right tool in front of the model the first time, the cheapest performance gain there is.

Security

3patterns

Prevents: An action nobody approved executes, and nobody can reconstruct it later.

Only allowlisted tools run, destructive calls wait for approve, edit, or reject, and every call leaves an audit record. Permission gating, Tool-call logs, Grounding.

RoadmapThe permission and audit middleware ships these three patterns as a layer other builders can adopt.

Reliability

6patterns

Prevents: Runs stall on a hung call, loop on a flaky one, or repeat a side effect after a resume.

Retries are classified and capped, every call runs on a clock with real cancellation, and re-runs are made safe by design. Bounded retries, Timeouts and cancellation, Idempotency.

RoadmapReference patterns shipped as code, so recovery behavior is inherited rather than rediscovered team by team.

Cost and observability

2patterns

Prevents: Token burn nobody can attribute, and agent behavior nobody can review.

Result budgets cap what re-enters the context on every step, and structured logs make per-tool cost and failure rates readable at a glance. Context budgeting, Tool-call logs.

RoadmapThe audit layer turns the log into the monthly cost and behavior record for every connected tool.

One call, start to finish

Anatomy of a tool call

Zoom into a single pass through the loop. Validation and the permission check sit before execution on purpose: a call stopped there costs nothing. The failure branch is designed, not discovered: errors drop to a bounded retry, and past the bound they escalate to a fallback or a person instead of looping.

Anatomy of one tool callA proposed call passes validation and a permission check, then executes. Errors go to a bounded retry or escalate to a fallback or a human. Successful results are shaped to a context budget before entering the model context.RequestValidationPermission checkExecutionResult shapingContextErrorbounded retryEscalatefallback or humanfailure branch: retry inside the bound,then escalate instead of loopingevery step writes one line to the audit log

And every step leaves a record

One line per tool call (AgentsUse shape)
{
  "ts": "2026-10-08T14:03:11Z",
  "run_id": "run_8f3",
  "step": 7,
  "tool": "web_search",
  "gate": "auto-approved",
  "args_hash": "b3:9f2c",
  "duration_ms": 1840,
  "outcome": "ok",
  "result_tokens": 3120,
  "attempt": 1
}

Arguments are hashed, not stored. Outcome, duration, and result size are what the monthly review reads. The pattern: Tool-call logs and observability.

Start from the failure

The failure you keep hitting, and the pattern that ends it

Nobody browses patterns for fun. Find the symptom your runs actually produce, then read the one pattern that addresses it.

Without harness work

Model wired straight to toolsWithout a harness, the model calls tools directly. Malformed arguments execute, a hung call stalls the run, a verbose result floods the context, and a side effect repeats when the run resumes.ModelToolsMalformed arguments executeHung call stalls the runVerbose result floods the contextSide effect repeats on resume

With the harness patterns

Model, harness, toolsWith the harness between model and tools, arguments are validated, permissions checked, results budgeted, retries bounded, and every call logged, so the four failure modes on the left have an owner.ModelHARNESSvalidate · gate · budgetretry · log · groundnine patterns, one layerToolsArguments checked before executionTimeouts and cancellation on every callResults trimmed to a budgetRe-runs made idempotent

Same model, same tools. The left wire is what most demos ship; the right wire is the same calls with an owner for each failure mode. The table below maps each symptom to its pattern.

Symptom in your runsPatternJob
A flaky call loops forever, or a hopeless call gets retried until the budget is goneError recovery and bounded retriesRecover
Malformed arguments execute, or the model burns turns guessing the format you wantedArgument validation before executionConstrain
One verbose result evicts the instructions the model still neededContext budgeting for tool resultsBudget
The agent sends, deletes, or spends before anyone saw the planPermission gating and human approvalGate
A retry or resume repeats a side effect: two charges, two emails, one edit applied twiceIdempotency and safe re-runsRecover
Nobody can answer "what did the agent do, in what order, and why" after something goes wrongTool-call logs and observabilityRecord
One hung call stalls the whole run, and a cancelled tool keeps writing after the user said stopTimeouts and cancellationRecover
Calls that depend on each other race in parallel, or everything runs serially and the task takes ten times longer than it shouldSequencing dependent tool callsCoordinate
The final answer asserts numbers and quotes that no tool ever returnedGrounding and citing tool resultsRecord

The patterns library

Nine patterns, each sourced

How the frameworks score on these patterns. Twelve frameworks and agent runtimes compared cell by cell on permissions, approvals, recovery, logs, and tool-result context, sourced from their official docs. Paperclip and Goose included, with their roles named.

Compare frameworks

The AgentsUse framework is in the works. Roadmap, not a release: a harness layer built from the patterns on this page. Builders who want in early can ask.

Request early access

Inside the directory

Harness notes, where the vendor docs support them

Five profiles now carry a harness-notes section: repeat behavior, result volume, scope rules, and polling behavior, each taken from the vendor's own documentation and dated. Profiles without a verifiable fact stay empty on purpose.

Tool profile

Harness notes: result volume is the budget lever

max_results accepts 1 to 20, and Tavily's agent guidance sets 5 as the starting point. Every result re-enters the context on each following step, so a high ceiling taxes the whole run, not just one call.

Open the profile

Tool profile

Harness notes: poll crawls to a terminal status

Crawl responses page results at 10 MB per response and hand back a next URL. The next URL also appears while pages are still processing, so treat job status (completed, failed, cancelled) as the end signal, never the absence of a next link.

Open the profile

MCP profile

Harness notes: scope the server before the model runs

The --read-only flag (or GITHUB_READ_ONLY=1) restricts the server to read tools, and read-only wins over --tools selections: a wrapper cannot widen what the flag closed. That is permission gating enforced by the server instead of the prompt.

Open the profile

MCP profile

Harness notes: snapshots are the context cost

The server works from structured accessibility snapshots rather than screenshots, which keeps each observation text-shaped and referenceable (click by ref) instead of image-priced.

Open the profile

MCP profile

Harness notes: scope is dynamic, repeat behavior is per tool

Allowed directories come from the launch arguments or from client Roots. When Roots are present they replace the configured list entirely, and they can change while the server runs. Scope checks belong in the harness on every call, not just at connect time. With no allowed directory from either source, the server fails at startup.

Open the profile

More notes as the facts land

A note appears only when a vendor document states the behavior. Browse the tools directory and the MCP directory for the verified setup, pricing, and limits record behind each profile.

Roadmap · UpcomingNot shipped, not dated, not priced

The AgentsUse framework

The directory maps the tool layer. This pillar teaches the runtime layer around it. The next step is AgentsUse shipping that layer as a framework of its own, built from the patterns above. Three directions, all pulled from gaps the catalog keeps hitting:

  • Compatibility truth: which tool works with which framework and model, answered from verified records instead of model memory.
  • Permission and audit middleware: allowlists, approval checkpoints for destructive calls, and one structured log line per tool call.
  • Reference harness patterns as code: validation, bounded retries, context budgets, and safe re-runs, shipped as components instead of advice.

This section is a roadmap statement. Nothing here is a shipped capability, and when something does ship it will be listed in the directory under the same verification rules as every other profile in it.

Mention the framework in your message. No date is promised; the early-access list hears first when there is something to run.

Questions

What is an agent harness?

The code and policy around a model that turns its intent into tool calls: it picks the tool, checks the arguments, enforces permissions, handles errors and timeouts, decides how much of each result stays in context, and logs what happened. The model proposes; the harness disposes.

Is the harness the same thing as the framework?

They overlap. A framework (LangGraph, CrewAI, the OpenAI Agents SDK) ships harness machinery such as retry policies, approval interrupts, and checkpointing. Your harness is the set of choices you make with that machinery: which tools run ungated, what the retry bound is, what gets logged. Two teams on the same framework run different harnesses.

Do I need all nine patterns?

No. Argument validation and a per-call timeout pay for themselves on day one. Add approval gates before the agent touches anything irreversible, logs before it runs unattended, and the rest as the failures show up. The checklist guide orders all nine by what breaks first.

Where does the directory fit in?

The harness assumes tools worth calling. The directory profiles verify setup paths, auth, pricing, and limits for the tools and MCP servers the harness drives, and five profiles now carry harness notes: vendor-documented facts about repeat behavior, result volume, and scope that a harness design needs.