Skip to main content
AgentsUse

Harness pattern · Record

Tool-call logs and observability

Log every call with its arguments (redacted), decision, duration, and result size, then read the logs as maintenance data, not just incident evidence.

The failure it prevents

In production: Nobody can answer "what did the agent do, in what order, and why" after something goes wrong.

An agent does something wrong at 14:03 and the investigation starts from a chat transcript and a shrug. Transcripts show what the model said, not what the harness did: which calls ran, which were gated, what the tool actually returned, how long it took, and what it cost. Without that record, every incident is unrepeatable and every argument about what happened is memory against memory.

The second failure is slower. Tool descriptions rot. A vendor changes a response shape, a parameter gains a new constraint, and the only symptom is a model that "got worse". Anthropic reads this directly off their evaluation traffic: total tool calls, redundant calls, errors, and invalid-parameter errors per tool tell them which descriptions to fix. If you do not count, you cannot see it.

How it works

Log at the harness, not in the model's narration. One structured record per call: timestamp, run id, tool, redacted arguments or an argument hash, the gate decision, duration, outcome class (ok, tool error, protocol error, cancelled), and result size in tokens. MCP gives servers a logging channel (notifications/message with syslog severity levels) and tells clients to keep configuration so users can enable, disable, and filter it; treat host-side logging of tool usage as an audit duty, not a debug extra.

Redact by default. Arguments and results carry secrets, personal data, and payloads you do not want in a log store. MCP's logging guidance is blunt: never emit secrets or credentials. Hash or truncate arguments, keep full text only where a retention decision says so.

Trace runs end to end. Anthropic runs full production tracing on their multi-agent system, watching decision patterns and tool usage while respecting privacy by monitoring patterns rather than reading conversation contents. Steal that line: aggregate first, drill into individual runs only during an investigation.

Turn counts into maintenance. Per tool: call volume, first-attempt failure rate, invalid-parameter rate, average result tokens, p95 duration. Review monthly. The tool with the rising invalid-parameter rate gets a better description; the tool returning 40,000 tokens per call gets a budget.

When to use it

  • Any agent running unattended, on a schedule, or against real accounts.
  • Approval-gated tools, where the decision record is the audit trail.
  • Multi-agent systems, where responsibility diffuses across workers by design.
  • Cost-sensitive deployments; result-token counts per tool find the spend.

When to skip or soften it

  • Local experiments in a sandbox with throwaway data; a console print is enough.
  • Extremely high-volume read calls where sampling beats completeness; sample deliberately and say so in the config.

Tradeoffs

Logs are a liability as well as an asset: they store what the agent touched, which is often the sensitive part. Retention limits, redaction, and access control are part of the pattern, not follow-up work. Volume is the other cost; one JSON line per call is cheap, full result bodies are not. Log sizes and hashes, store bodies separately under a shorter retention.

One line per tool call (AgentsUse shape)

The minimum record that answers the incident questions and feeds per-tool maintenance metrics. Arguments are hashed; result bodies live elsewhere under shorter retention.

One line per tool call (AgentsUse shape)
{
  "ts": "2026-10-08T14:03:11Z",
  "run_id": "run_8f3",
  "step": 7,
  "tool": "tavily_search",
  "gate": "auto-approved",
  "args_hash": "b3:9f2c",
  "duration_ms": 1840,
  "outcome": "ok",
  "result_tokens": 3120,
  "attempt": 1
}

Shape source: AgentsUse suggested shape; logging duties per MCP specification 2025-06-18. Placeholders only; adapt names and numbers to your stack.

Implementation checklist

  • Every tool call emits one structured record, including gated, rejected, and cancelled calls.
  • Arguments are redacted or hashed; secrets never reach the log (MCP logging rule).
  • Outcome classes separate ok, tool errors (isError), protocol errors, timeouts, and cancellations.
  • Result tokens and duration are recorded per call and aggregated per tool monthly.
  • Retention and access rules for logs are written down next to the logging config.
  • Invalid-parameter and first-attempt failure rates per tool feed the description-fix list.

Sources and freshness

Claims on this page checked against these sources on 2026-10-08. Code and config blocks are shapes to adapt, not benchmarks.