Harness pattern · Gate
Permission gating and human approval
Sort tools by consequence, let the safe ones run, and hold the irreversible ones for a decision: approve, edit, reject, or answer.
The failure it prevents
In production: The agent sends, deletes, or spends before anyone saw the plan.
The failure here is not a wrong answer. It is a correct-looking action nobody authorized: the email sent to the wrong list, the rows deleted instead of archived, the purchase made twice. It happens in one step, at machine speed, and undo ranges from awkward to impossible.
Tools advertise their own risk. MCP tool annotations carry readOnlyHint, destructiveHint, and idempotentHint, and the filesystem server sets them per tool: edit_file is neither idempotent nor non-destructive, create_directory is idempotent. A harness that ignores those hints is declining free information.
How it works
Gate by consequence, not by tool count. Read-only tools run freely. Writes pause based on what they change. The gating decision can depend on the arguments, not just the tool name: LangChain's HumanInTheLoopMiddleware accepts a `when` predicate that inspects the actual tool call, so reading any file is fine but writing outside the project directory pauses.
Offer four decisions, not two. Approve runs the call as proposed. Edit runs a human-modified version. Reject blocks it and tells the model why. Respond answers the agent with information instead of executing. LangChain ships exactly this decision set, and the reject path matters: the reason goes back into the context, so the model proposes something else instead of retrying the same call.
Fail closed. The OpenAI Agents SDK does this by default in its approval flow: when tool arguments are malformed, the tool is not invoked and goes to manual approval. Anything the harness cannot parse goes to a human, never to the tool.
MCP writes the host-side duties into the specification: hosts should show tool inputs to the user before calling tools, confirm before sensitive operations, and validate results before passing them on. Elicitation covers the structured-question case, with accept, decline, and cancel as first-class answers the harness must handle.
When to use it
- Any tool that sends, deletes, publishes, purchases, or changes permissions.
- Writes whose blast radius depends on arguments (paths, recipients, amounts, row counts).
- Tools annotated destructiveHint: true; the vendor already told you.
- First deployments of a new tool pairing, loosened only after a run of clean reviews.
When to skip or soften it
- Read-only tools with scoped credentials; approval theater on reads trains people to click approve blindly.
- Reversible sandboxed writes (drafts, staging data) where undo is one click.
- Batch jobs where each item was approved as a class beforehand and logged individually.
Tradeoffs
Every gate is a delay and a chance the run stalls waiting for a person who went to lunch. Gate too much and the agent is a very slow form. The calibration is consequence times irreversibility: a draft email needs no gate, sending it does, and editing 4,000 rows needs a gate even if each row is small. Keep the decision log; the record of who approved what is most of the value.
Approval gates per tool (LangChain middleware)
LangChain's human-in-the-loop middleware pauses on configured tools and resumes with the human decision. The vocabulary below (interrupt_on, allowed_decisions) is from the LangChain documentation; the tool names are placeholders.
from langchain.agents.middleware import HumanInTheLoopMiddleware
HumanInTheLoopMiddleware(
interrupt_on={
"read_data": False, # auto-approve
"send_message": { # human decides
"allowed_decisions": ["approve", "edit", "reject"],
},
"delete_records": { # never editable, approve or reject
"allowed_decisions": ["approve", "reject"],
},
},
description_prefix="Approval needed",
)Shape source: LangChain documentation: Human-in-the-loop middleware. Placeholders only; adapt names and numbers to your stack.
Implementation checklist
- Every tool is classified: auto-run, gate, or never-autonomous, and the classification is written down.
- Vendor hints (readOnlyHint, destructiveHint) feed the classification instead of being rediscovered by hand.
- Gates can inspect arguments, so a safe call to a powerful tool still runs.
- Reject and edit return reasons and edits to the model; a bare "no" teaches nothing.
- Malformed or unparseable arguments go to a human by default, never to the tool.
- Approvals, rejections, and edits are written to the tool-call log with the deciding identity.
Sources and freshness
- LangChain documentation: Human-in-the-loop
- MCP specification 2025-06-18: tools (annotations, host duties)
- MCP specification 2025-06-18: elicitation (accept, decline, cancel)
- OpenAI Agents SDK: human-in-the-loop and approvals
Claims on this page checked against these sources on 2026-10-08. Code and config blocks are shapes to adapt, not benchmarks.