Guide
The Production Harness Checklist: 24 Checks Before an Agent Touches Real Work
Twenty-four pass-or-fail checks across the call, the gate, the run, and go-live, drawn from Anthropic's engineering posts, the MCP specification, and LangGraph and OpenAI documentation. Print it, score it, and widen scope only where every box holds.
2026-10-0814 min read
A demo agent needs a model and some tools. A production agent needs a harness that survives contact with real systems: flaky vendors, fat results, impatient users, and its own pauses. The difference rarely shows in week one. It shows at the first 2 a.m. incident, priced in duplicate sends and unanswerable questions.
This checklist is the audit we would want run against any agent before it touches real work. Twenty-four checks, each pass-or-fail on evidence, grouped around the call, the gate, the run, and the moment you widen scope. Every check traces to a published source: Anthropic's engineering posts on tools, context, and multi-agent systems; the MCP specification (version 2025-06-18); and the LangGraph and OpenAI Agents SDK documentation. The pattern pages in this pillar carry the mechanisms and the links. This page is the scoring sheet.
How to use it: one pass, in order, with the logs open. For each check, produce the artifact (the config value, the log line, the recorded test) or mark it failed. Fixes go to the failed checks in order; the groups are sequenced so that early failures are the cheap ones.
One rule keeps the exercise strict: evidence beats testimony. 'The agent basically retries sensibly' is not check 5. Check 5 is a policy value and a log line showing a permanent error that was not retried. If the team cannot produce the artifact in five minutes, the check fails, however good the intentions are. This sounds harsh and saves real time, because it converts architecture arguments into a scavenger hunt with a definite end.
Around every call: checks 1 to 8
These eight checks govern a single tool call from proposal to result. They are the cheapest in the list and remove the largest share of everyday failures.
- Check 1, schema validation before dispatch: arguments are checked against the tool's published inputSchema in the harness, not discovered as errors afterward. MCP puts validation duties on both server and client; doing it client-side catches the error while it is still cheap. Evidence: the validation step in the call path, plus rejection counts per tool.
- Check 2, rejections that teach: a failed validation returns the field, the expected shape, and one valid example. Anthropic's guidance is explicit that error responses should give the model specific, actionable improvements, not opaque codes. Evidence: three real rejection messages from your logs, readable by a new teammate.
- Check 3, a written timeout per call: every tool has a maximum call duration sized from measured latency, and the run has a hard ceiling behind it. The MCP specification advises timeouts on all requests; unmeasured guesses are acceptable on day one and get replaced with p99-based numbers from check 18. Evidence: the per-tool timeout table.
- Check 4, cancellation that lands: on timeout or user stop, the harness sends cancellation (notifications/cancelled in MCP, with request id and reason), stops waiting, and drops late responses instead of applying them. Evidence: a log line showing a cancelled request and a late response being ignored.
- Check 5, bounded, classified retries: retries apply to transient failures only, never to validation or programming errors, with attempt caps and exponential backoff with jitter. LangGraph's published defaults (3 attempts, 0.5s initial, doubling, 128s cap) are a defensible starting row in the table. Evidence: the retry policy, and a log showing a permanent error not retried.
- Check 6, a recorded repeat class per write tool: idempotent, idempotency-keyed, check-then-write, or never-repeat, decided from vendor hints (MCP's idempotentHint) or a controlled double-run. The filesystem server's documented spread (create_directory idempotent, edit_file not) shows why one policy cannot cover every tool. Evidence: the classification table next to the retry policy.
- Check 7, a result-size budget: each tool's output is capped (Claude Code's default is 25,000 tokens per tool response, per Anthropic), trimmed results say so and name the way to get more, and search-style tools default to low result counts. Evidence: the budget table and one trimmed result from the logs.
- Check 8, results validated before use: tool output passes a shape check (the declared outputSchema where one exists) before it enters the context, and malformed output becomes a tool error, never citable data. MCP assigns result validation to the host explicitly. Evidence: one malformed result caught in a test, with the error the model received.
Gates and permissions: checks 9 to 13
Five checks decide who or what approves the calls that can hurt someone. If your agent has no write tools, record that fact here and move on; the checks return the day it gets one.
Two failure stories explain why this group exists. In the first, an agent with broad mail access sends a well-crafted message to a real list at 3 a.m., because nothing in the runtime distinguished draft from send. In the second, a human approver, worn down by forty trivial prompts a day, waves through the one call that mattered. Both are gating failures: the first gate missing, the second gate exhausted. The checks below calibrate for both, consequence on one axis and human attention on the other, because a gate nobody reads is a formality, not a control.
- Check 9, a written consequence classification: every tool is auto-run, gated, or never-autonomous, and the list lives in config, not in the system prompt. Prompts are advice; config is enforcement. Evidence: the classification table, reviewed when the tool set changes.
- Check 10, argument-aware gates: the gate can inspect the call's arguments, so reading any file runs free while writing outside the project pauses. LangChain's middleware supports exactly this with a predicate on the tool call. Evidence: two logged calls to the same tool, one auto-approved and one gated, differing by arguments.
- Check 11, a full decision vocabulary: approve, edit, reject, respond. Reject carries a reason back to the model; edit runs the corrected version. Yes/no gates starve the model of the information it needs to propose better. Evidence: one rejected call in the logs with its reason, and the model's next proposal.
- Check 12, fail-closed on the unparseable: arguments the harness cannot parse never reach the tool; they go to a human. The OpenAI Agents SDK implements this default in its approval flow. Evidence: a deliberately malformed call in testing, routed to approval, with the tool never invoked.
- Check 13, narrowest-scope credentials, enforced server-side where possible: read-only flags and scoped tokens beat prompt instructions. The GitHub MCP server's --read-only flag overrides tool selections, so the ceiling holds even if the wrapper is misconfigured. Evidence: the scope list per credential, and one test call refused by scope, not by politeness.
Around the run: checks 14 to 19
Six checks govern the run as a whole: what it remembers, what it records, and how anyone knows it is drifting.
Single calls forgive sloppy state. Runs do not. A run is where a mid-task crash meets yesterday's half-finished work, where a result fetched in step two gets cited in step forty, and where the costs nobody itemized quietly compound. Production pain concentrates here for a structural reason: demos exercise calls, production exercises sequences, pauses, resumes, and partial failures. These six checks are written against the sequence, and most first-time failures of the whole checklist live in this group.
- Check 14, one structured record per call: timestamp, run id, tool, decision, redacted arguments or a hash, duration, outcome class, result size, attempt. Secrets never reach the log; MCP's logging guidance prohibits emitting credentials through log channels. Evidence: the record for the last real incident, answering what ran, in what order, with what outcome, without the transcript.
- Check 15, durable step completion: when a step finishes, that fact is stored before the next step starts, so a resume skips finished work instead of replaying it. LangGraph's checkpointing with pending writes is the reference behavior. Evidence: the kill-and-resume test from check 20, passing.
- Check 16, dependencies as data: any call consuming another call's output is an edge in a written step graph. Parallel execution is reserved for calls that consume nothing from each other. Evidence: the graph for the longest production task, with at least one gate between a read and the write that trusts it.
- Check 17, a compaction rule for long runs: past a stated size, the run summarizes state (decisions, open issues, current artifacts) and clears raw results instead of dragging them. Anthropic's context guidance lists tool-result clearing as the lightest compaction and describes summaries plus the most recent files surviving. Evidence: a long run whose context stayed bounded, with the summary it produced.
- Check 18, a monthly per-tool metrics review: invalid-parameter rate, first-attempt failure rate, average and p95 result tokens, p95 duration, per tool, trended. These are the signals Anthropic publishes for its own agent evaluations. Evidence: last month's table and the one change it caused (a description rewritten, a budget tightened, a tool replaced).
- Check 19, a fixed evaluation set: ten to twenty realistic tasks, run after any harness, tool, or model change, with failures read. Anthropic reports that a set of about twenty queries surfaced real regressions in their multi-agent work. Evidence: the set, its last run date, and its last caught regression.
Before you widen scope: checks 20 to 24
The last five checks run at the moment of temptation: the agent has behaved for a month and someone wants to remove a gate or add a tool. These are the graduation exams.
They are grouped separately for a reason. The first nineteen checks describe the system as built; these five describe the system as stressed. Tests 20 to 22 deliberately break things (kill the process, run the write twice, feed a stale result) because the failures that matter in production are precisely the ones polite testing never meets. Checks 23 and 24 are organizational, and they are the ones skipped most often, on the theory that the team will remember to be careful. Teams do not remember. Writing wins.
- Check 20, the kill-and-resume test: start a real task, kill the process mid-step, resume, and audit every side effect. Duplicates fail the test and send you back to checks 6 and 15. Evidence: the test record, with effect counts before and after.
- Check 21, the double-run test per write tool: execute each write tool twice with identical inputs in a sandbox and record what the second run does. This settles repeat classes empirically where vendor hints are absent. Evidence: the results table feeding check 6.
- Check 22, the stale-result test: feed the run a cached, outdated, or post-cancellation result and confirm the harness detects it (validation failure, provenance mismatch, ignored late response) rather than citing it. Evidence: the test and the error or drop it produced.
- Check 23, a written exit rule: before scope widens, name the condition that pulls the task back to manual (an error class, a cost ceiling, a drift metric from check 18). Deciding in advance is what makes widening calm. Evidence: the rule, in the runbook, with the metric it watches.
- Check 24, a review cadence with a name on it: someone specific reads the logs on a schedule, and a named symptom (rising invalid-parameter rate, first duplicate effect) tightens gates automatically. Approval records are only an audit trail if somebody audits. Evidence: the last review's date and one action it produced.
Scoring and what to do with failures
A production-ready harness passes all twenty-four, or carries a written exception with a reason and a review date. In practice, first audits cluster: teams pass the call-level checks, fail durability (6, 15, 20), and discover their logs cannot answer check 14's question. That ordering is good news. The failures are concentrated in the run-level group, and they are buildable in days, not quarters.
Treat scope as the reward for boring logs. When the monthly review shows flat failure rates and the evaluation set is green, widen one notch: one more auto-approved tool, one higher budget, one longer unattended window. When a metric moves, tighten back to the last boring configuration and diagnose by layer. This ratchet, widen on evidence and narrow on signal, is the operating loop behind every pattern in this pillar. The checklist is how you know which notch you are standing on.
A worked scoring makes this concrete. Picture a customer-support agent with read access to the help center, draft access to replies, and a gated send. First audit: checks 1 to 5 pass (validation and timeouts came free with the framework, once configured). Check 6 fails: nobody knows whether the draft-save tool is safe to repeat. Check 11 fails: the gate offers approve and deny, and denied drafts teach the model nothing. Checks 14 and 15 fail together: logs show messages, not calls, and a restarted run redrafts from scratch. That is 17 of 24, and the fix list writes itself: record the draft tool's repeat class, add edit and reject-with-reason to the gate, log calls properly, persist draft completion. Two afternoons later the same audit scores 23, with check 22 scheduled. This is what progress looks like in this discipline: not a smarter agent, a shorter list of named ways it can still hurt you.
Who holds the pencil matters. Self-audits drift generous within two cycles; the person who wrote the retry policy reads check 5 and remembers the intent, not the evidence. Rotate the auditor, or pair the audit with the monthly metrics review so the numbers argue with the scores. The checklist is also the right artifact to hand a new teammate: twenty-four questions whose answers describe the system more truthfully than any architecture diagram, because each answer points at a running mechanism rather than a drawn box.
Keep the completed sheets. A run of dated checklists is the harness's medical record: when the score moved, which change moved it, and what broke the last time a gate was loosened. When an incident happens anyway (one will), the last clean sheet shortens the investigation by telling you what the system looked like when it was healthy, and which check to re-examine first. Production readiness is not a state you reach. It is a score you keep proving.
Exceptions deserve the same rigor as passes. When a check cannot apply yet (the agent reads one internal wiki and nothing else, so the go-live group waits), write the exception beside the check with the condition that retires it: first external tool, first write, first unattended schedule. Undated, unconditional exceptions are how checklists rot into wall art. Dated ones become the roadmap for the next widening, which is the entire point of keeping score.
What usually fails first, and the order to fix it
First audits tend to rhyme. The most common failure is check 14, the log that cannot answer questions: teams log model messages, or tool names without outcomes, and discover during the first incident that they cannot reconstruct the run. Fix it first regardless of group order, because every other check's evidence lives in that record. The second is the durability pair, checks 6 and 15, usually failed together: nobody wrote down what a re-run does, because the demo never crashed. The kill-and-resume test (check 20) exists to make that gap visible while the stakes are still a sandbox.
The third common failure is quieter: check 18, the monthly review that never happens, so budgets and timeouts freeze at their day-one guesses while vendors change shapes and prices underneath. The symptoms arrive slowly (rising result tokens, a creeping invalid-parameter rate) and get attributed to the model. Fix order after the audit: logs first, durability second, gates third (they are fast and make unattended running defensible), metrics cadence fourth, and budgets once the metrics exist to size them. Context failures feel urgent and diagnose badly without data, so resist fixing them first; the order above is the difference between a checklist and a wish list.
A final caution about the gate group. Passing checks 9 to 13 on paper while a human rubber-stamps forty approvals a day is a failed gate wearing a passed costume. The checklist cannot see approval quality, so measure it: median time-to-decision on gated calls, and the share of approvals where the arguments were actually displayed. If attention is being spent on trivia, reclassify until the queue is short enough to read. The pillar's diagnosis guide works through that failure, the gate that cried wolf, in full.
Continue in the Harness pillar
The pattern pages behind this guide: the failure each one prevents, how it works, and the checklist for shipping it.
Frequently asked questions
How is this different from the pre-flight checklist guide already on this site?
The pre-flight checklist decides whether a task should go to an agent at all and how to test the choice. This checklist assumes the task was chosen and audits the runtime around it: validation, timeouts, gates, logs, resume behavior. Run the first when selecting work, the second before widening what an approved agent may do unsupervised.
Do I score partial credit on a check?
No. Each check is written to be pass or fail on evidence: show the timeout value, the log line, the recorded repeat class. A check you almost pass is the one that fails at 2 a.m. Where a check genuinely does not apply (an agent with no write tools skips the double-run test), write the reason next to it instead of leaving it blank.
How often should I rerun the checklist?
On three triggers: a new tool joins the set, a tool or model version changes, and any production incident. Between triggers, the monthly metrics review (check 18) watches for drift, rising invalid-parameter rates or result sizes that no single run would reveal.
Which checks do teams fail most?
In our reading of published postmortems and framework documentation, the repeat offenders are checks 6 and 15 (nobody recorded what a re-run does), check 14 (logs that cannot answer an incident question), and check 23 (no written exit rule, so scope only ever widens). All four are cheap on day one and expensive to retrofit.
Keep reading
Guide
Agent Harness Engineering: Where Tool Use Succeeds or Fails
The model gets the credit and the blame, but the harness decides what actually happens: which call runs, with what arguments, under whose permission, and what happens when it fails. A sourced guide to the layer, its six jobs, and the order to build them in.
2026-10-08
Guide
Diagnosing Tool-Call Failures: Read the Run, Fix the Layer
An agent run failed. The transcript will not tell you why. A field method for separating argument errors, tool errors, timeouts, context bloat, permission blocks, and sequencing races, with the fix that belongs to each layer.
2026-10-08
Guide
A Checklist Before Letting an AI Agent Touch Real Work
Twelve checks to run through before an AI agent handles anything that matters - customers, money, or your reputation.
2026-09-26