Skip to main content
AgentsUse

Harness pattern · Budget

Context budgeting for tool results

Treat tool output as spending from a fixed budget: cap it, shape it, keep the recent full text, and clear what the run has already used.

The failure it prevents

In production: One verbose result evicts the instructions the model still needed.

Every token a tool returns stays in the conversation and competes with the instructions, the task, and every result after it. Anthropic names the effect context rot: as the token count grows, the model's ability to recall and use what is in the window degrades. A single 200 KB search result or accessibility tree does not just cost money. It pushes the original instructions toward the part of the context the model uses worst.

The failure looks like forgetfulness, not like an error. The agent repeats a step, drops a constraint from the brief, or re-reads a page it already read. Nothing crashed, so teams blame the model and swap it, when the fix was to stop flooding its desk.

How it works

Cap every result. Anthropic restricts tool responses in Claude Code to 25,000 tokens by default. Set your own cap per tool, enforced in the harness, not requested in a prompt.

Let the caller choose the shape. Anthropic documents a response_format enum (concise versus detailed) where the detailed Slack response ran 206 tokens and the concise one 72, about a third, with detail reserved for the calls that need identifiers for downstream steps. Pagination, range selection, filtering, and truncation with sensible defaults are the same idea at larger scale: return the slice, plus a way to ask for more.

Clear and compact as the run ages. Once a tool result deep in the history has been used, the raw text rarely needs to stay; Anthropic lists tool result clearing as the lightest form of compaction. For long runs, compact properly: summarize the conversation state, keep the decisions and open issues, drop redundant outputs, and continue in a fresh window. Their Claude Code compaction keeps the summary plus the five most recently accessed files.

Hold references, not payloads. Just-in-time retrieval keeps lightweight identifiers (file paths, queries, URLs) in context and loads the data when a step actually needs it. The model works from a map instead of a warehouse.

When to use it

  • Search, scrape, database, and browser tools, where result size is unbounded and caller-controlled.
  • Long runs with many calls, where results accumulate against a fixed window.
  • Any tool whose full output is usually scanned for one fact.
  • Multi-agent setups, where a worker can return a distilled summary instead of its raw exploration.

When to skip or soften it

  • Short runs with a handful of small results; the budget machinery costs more than it saves.
  • Results the model must quote verbatim later (legal text, exact figures); keep those whole and budget elsewhere.

Tradeoffs

Truncation hides the one line that mattered, and summarization can drop the qualifier that made a fact safe to use. Budget with an escape hatch: every trimmed result should say it was trimmed and how to get the rest, and steering text at the cut point ("refine the query for the remaining items") teaches the model cheaper behavior. The second tradeoff is bookkeeping. Budgets, clearing, and compaction are state your harness must maintain correctly across resumes; a bug there deletes evidence the model needed.

Harness result budget (AgentsUse shape)

A per-tool budget enforced at the harness, combining Anthropic's published defaults (25,000-token cap, concise/detailed response shapes) with retention rules. Adapt the numbers; keep the enforcement.

Harness result budget (AgentsUse shape)
# result budget, enforced by the harness (not requested in a prompt)
TOOL_RESULT_MAX_TOKENS=25000   # hard cap per result (Claude Code default, per Anthropic)
TOOL_RESULT_KEEP_FULL=3        # most recent results stay verbatim
TOOL_RESULT_OLDER=clear        # older results: clear, summarize, or file-reference
SEARCH_MAX_RESULTS=5           # caller-side ceiling for search tools
INCLUDE_RAW_CONTENT=false      # raw page bodies only on request

Shape source: Anthropic: Writing effective tools for AI agents; Effective context engineering. Placeholders only; adapt names and numbers to your stack.

Implementation checklist

  • Every tool has a maximum result size enforced in the harness.
  • Search and list tools paginate or accept a result ceiling; the harness sets a low default.
  • Trimmed results state that they were trimmed and how to retrieve the rest.
  • Results older than a few steps are cleared, summarized, or replaced by a reference.
  • Long runs compact: decisions and open issues survive; raw outputs do not.
  • Token use per tool is logged, so the expensive tools are visible by name.

Sources and freshness

Claims on this page checked against these sources on 2026-10-08. Code and config blocks are shapes to adapt, not benchmarks.