Writing SOPs Your Agents Can Follow
Agents follow procedures better when procedures are written for agents. How agent SOPs differ from human ones, and the anatomy of a procedure an agent can actually execute.
Published 2026-09-29 · 10 min read
Standard operating procedures are how organizations bottle expertise — and they are also how agents become reliable. An agent with a good SOP for a task performs dramatically better than one improvising from a vague instruction. But most SOPs are written for humans, full of assumed background knowledge, judgment calls, and steps like use your discretion that an agent cannot execute safely. Writing SOPs for agents is a distinct skill: same goal as human SOPs, different assumptions throughout. This guide covers how agent SOPs differ, what makes a procedure agent-followable, and how to test and maintain them.
How agent SOPs differ from human SOPs
Human SOPs lean heavily on shared context: the reader knows the systems, the jargon, the unwritten rules, and when to break the written ones. Agents have none of that. Every assumption a human SOP makes is a potential failure point for an agent — the step that says check the dashboard assumes the agent knows which dashboard, with what credentials, looking for what. Judgment calls are the sharpest difference: use your best judgment is executable by a person and meaningless to an agent unless the criteria for judgment are spelled out. Agent SOPs must also be explicit about scope boundaries — what is in bounds versus what requires escalation — because agents do not feel the social cues that tell humans they are oversteptpping. The result reads as more pedantic than a human SOP. That pedantry is the feature.
Anatomy of an agent-followable SOP
- Objective in one paragraph. What success looks like, stated concretely enough that the agent can recognize it — not process description, but outcome description.
- Preconditions. What must be true before starting: required access, required inputs, required tools. The agent checks these first instead of discovering gaps mid-run.
- Atomic steps. Each step is one action with a verifiable result. If a step needs a paragraph to explain, it is two steps.
- Decision points with criteria. Wherever the procedure branches, spell out the exact conditions for each branch — if X, do A; if Y, do B — with no appeal to judgment.
- Output specification. Exactly what the agent must produce: format, fields, length, destination. The handoff contract for the next consumer.
- Escalation triggers. The specific situations where the agent stops and asks a human — stated as observable conditions, not feelings.
- Worked example. One concrete walkthrough of the procedure on a realistic input, showing what each step's output looks like.
Writing steps agents can execute
The unit of an agent SOP is the verifiable step: an action plus a way to confirm it worked. Fetch the customer record and confirm the account status field reads active is a good step; look into the customer's situation is not — it has no completion criterion. Prefer imperative verbs tied to tools the agent has: query, read, write, send, check. Quantify where humans would use feel: review the last ten tickets, not review recent tickets. Name the systems explicitly: in the billing dashboard at this URL, not in the dashboard. And keep each SOP to one procedure — the urge to cover every variation produces documents agents cannot follow. Variations become separate SOPs or explicit branches, never footnotes the agent must interpret.
Testing your SOP
- Dry run on a realistic case. Have the agent execute the SOP on a real input while you watch — the gaps reveal themselves within minutes.
- Adversarial read. Hand the SOP to someone unfamiliar with the task and ask where they would get stuck; agents get stuck in the same places.
- Edge-case walkthrough. Run through the three weirdest real cases you can remember and check the SOP handles each without improvisation.
- Deliberate ambiguity test. Give the agent an input that falls between two branches and see whether it escalates correctly or guesses.
- Time it. If the SOP takes the agent five times longer than the human process, something in the procedure is fighting the agent's strengths.
Maintaining SOPs
SOPs rot. Systems change, tools get replaced, edge cases accumulate, and the procedure that worked in March silently breaks in September. Assign ownership: every SOP has a named owner responsible for its accuracy. Review on a schedule tied to change rate — monthly for fast-moving domains, quarterly for stable ones. Treat agent failures as SOP bug reports: when the agent fails at a step, the first question is whether the SOP was wrong, not whether the agent was. Version your SOPs and keep a changelog; when behavior changes, you want to know which edit caused it. And prune ruthlessly — an SOP library with forty stale procedures is worse than ten current ones, because agents cannot tell which are stale.
Writing for agents teaches you to write better procedures for everyone: the explicitness, the verifiable steps, the defined escalation paths — all of it transfers. Start with your single most repeated agent task, write the SOP to the standard above, dry-run it, and watch reliability jump. Then do the next one. Boring, systematic, and extraordinarily effective.
Versioning SOPs like code
SOPs change, and unversioned changes create mystery failures: the agent's behavior shifts and nobody knows which edit caused it. Keep SOPs in version control with the same discipline as code. Every change gets a commit message explaining why — not what changed (the diff shows that) but the reason: which failure or observation motivated it. Tag versions that correspond to known-good behavior, so you can roll back when a new version degrades results. Review SOP changes like code changes: a second pair of eyes catches the ambiguous step you wrote at midnight. And keep a changelog the agent itself can read — when the agent's behavior changes after an SOP update, the first debugging step is reading what changed, and a good changelog makes that a thirty-second task instead of an archaeology project.
- Commit messages explain why the SOP changed, not just what — the reason is what future debuggers need.
- Tag known-good versions so rollback is one command, not a reconstruction.
- Review SOP edits like code edits: ambiguity is a bug, and fresh eyes catch it.
- Keep a human-readable changelog the agent can consult when its behavior shifts.
- Never edit the production SOP directly — change, test, then promote, exactly like code.
SOPs for failure modes, not just happy paths
Most SOPs describe the sunny-day workflow and leave the agent improvising when things go wrong — which is precisely when improvisation is most dangerous. Every SOP needs a failure section: what to do when the API returns an error, when the expected file is missing, when the data looks wrong, when a step takes ten times longer than usual. For each failure mode, specify the detection (how the agent recognizes it), the response (retry, skip, escalate, halt), and the reporting (what gets logged and who gets told). These sections feel pessimistic to write and are the most-used parts of the document in production. An SOP without failure handling is a demo script; an SOP with failure handling is an operations document. Write for the bad days, because the good days need no instructions.
When the agent should deviate from the SOP
Blind SOP compliance is its own failure mode: the agent follows steps that are clearly wrong for the situation because the document said so. Build in explicit deviation authority with boundaries. The SOP should state which steps are inviolable (safety checks, approval gates, audit logging) and which are guidelines the agent may adapt when the situation warrants. Require the agent to log deviations with reasons — this creates accountability and, more usefully, a record of where the SOP was wrong. Review deviation logs regularly: frequent deviation from the same step means the step is broken, not the agent. The goal is a living document improved by field experience, not a rigid script that punishes judgment. The best SOPs get shorter over time as the inviolable core distills out of the accumulated deviations.
The quarterly SOP audit
SOPs decay: tools change, edge cases accumulate, workarounds fossilize into steps nobody remembers the reason for. Audit quarterly with fresh eyes — ideally someone who did not write the SOP. The auditor runs the SOP against current reality: does each step still work, is each tool still available, do the failure sections match the failures actually occurring? They also read the deviation and incident logs for patterns the daily operators stopped noticing. Common findings: steps that reference deprecated tools, approval gates for risks that no longer exist, missing coverage for the failure mode that caused last month's incident. The audit takes a few hours and prevents the slow rot that turns a good SOP into a cargo cult — steps followed because they are written, not because they work.
SOPs as onboarding for humans too
A well-written agent SOP is also the best onboarding document for the humans who will supervise the agent. New team members reading the SOP learn the workflow's logic, its failure modes, and its boundaries — the same knowledge they need to oversee the automation competently. This dual use imposes a useful discipline: if a competent new hire cannot understand the SOP, the agent probably cannot follow it reliably either, because both are intelligent readers working from the document alone. Write with both audiences in mind: precise enough for the agent's literal execution, clear enough for a human's first week. Keep a human-oriented preamble — why this workflow exists, what good looks like, who to ask — that the agent can skip but the newcomer needs. Teams that maintain SOPs as shared human-agent documents get two benefits for one maintenance cost, and the humans supervising agents actually understand what they are supervising.
- If a new hire cannot follow the SOP, the agent probably cannot either — use humans as the readability test.
- Precise enough for literal execution, clear enough for a first-week employee.
- A human preamble — purpose, success criteria, who to ask — that the agent skips and the newcomer needs.
- Shared documents mean supervisors actually understand the workflows they oversee.
- One maintenance cost, two audiences: the economics favor dual-use documentation.
Keep reading
A practical look at where AI agents genuinely help in 2026 and where they still fall short, so you can delegate with confidence.
How to Write Instructions AI Agents ActuallyClear, practical techniques for writing agent instructions that get followed — with examples of vague versus precise wording.
AI Agent vs. Chatbot vs. Script: What's theAgents, chatbots, and scripts get lumped together. Here is what each one actually is, and when to use which.