Multi-Agent Workflows That Actually Work

When teams of AI agents help, when they hurt, and the four collaboration patterns that earn their keep — plus the contracts and controls that make them work.

Published 2026-09-29 · 10 min read

The pitch is seductive: instead of one AI agent, deploy a team of them — a researcher, a writer, a critic, a coordinator — and watch complex work assemble itself. The reality most people encounter is less cinematic: agents talking past each other, duplicating work, burning through budget while producing less than a single well-instructed agent would have. Multi-agent setups are not useless, but they are a specialized tool, and most tasks do not need them. This guide covers when multiple agents genuinely help, the patterns that earn their keep, and how to avoid the coordination failures that sink most attempts.

When more agents actually help

  • Genuinely parallel subtasks. Researching several competitors, where each lookup is independent, is the canonical win: fan the work out, collect the results, done.
  • Generator-critic loops. One agent drafts, a second reviews against explicit criteria. The critic needs no tools and a narrow brief — it is cheap and catches real errors.
  • Specialist pipelines. Sequential handoffs where each stage has a distinct skill: extract data, then analyze it, then write it up. Each agent stays in its lane.
  • High-stakes judgment calls. Two agents arguing from different briefs can surface risks a single agent glosses over — useful for decisions with real downside, wasteful for routine work.
  • Long-running background work. A monitor agent watching for events while you do other things is multi-agent in the trivial sense, and it works because the agents barely interact.

Why most multi-agent setups fail

The failure modes are consistent enough to list. Coordination overhead: agents spend more context negotiating with each other than doing the work. Duplicated context: every agent needs the background, so you pay for the same briefing several times over. No shared state: each agent keeps its own notes, they diverge, and nobody notices until the final assembly contradicts itself. Circular deliberation: without a decider, agents debate past the point of usefulness because disagreeing is cheaper than committing. And cost multiplication, the quiet killer: several agents each making tool calls turns a reasonable run into an expensive one, often for results a single agent would have matched. If your multi-agent setup feels like managing people, that is because you have recreated the hard parts of management without the parts that make human teams work.

Contracts between agents

Agents that hand work to each other need contracts as explicit as API schemas. Define exactly what the upstream agent delivers: format, fields, length limits, and what done means. Research the competitor is not a contract; return a markdown table with columns for pricing, key features, and limitations, maximum 30 rows, with a source link per row is. The downstream agent's prompt should restate what it expects to receive and what to do when the input is malformed — ask for a fix, do not silently improvise. Write these contracts down where you can see them, because when a pipeline produces garbage, the contract is the first place to look: either it was violated or it was too vague to be violated.

Shared state without chaos

Give the team one shared state file and strict rules about who writes what. The coordinator owns the plan and the status; workers own their own findings sections and touch nothing else. Timestamps and owner labels on every entry prevent the classic failure where two agents update the same section with contradictory conclusions. Workers should write results, not commentary — the state file is a database, not a chat log. And designate a single decider for conflicts: when worker A says the data supports X and worker B says it supports Y, someone has to rule. That someone can be you, or a coordinator agent with an explicit tie-breaking instruction, but it cannot be nobody.

Keeping costs under control

  • Cap the fan-out. Three parallel workers is usually plenty; ten is usually theater. Each additional worker adds coordination cost faster than it adds throughput.
  • Match models to roles. Workers doing extraction or formatting run fine on smaller, cheaper models; reserve the strongest model for the coordinator and the critic.
  • Set a team budget, not just per-agent limits. Multi-agent runs fail expensively; a global cap with an alert at seventy percent prevents surprises.
  • Kill switches for loops. If worker outputs stop changing between rounds, stop the run — the agents have converged or stalled, and further rounds add nothing.
  • Reuse briefings. Write the shared background once, store it in a file, and have workers read it instead of each receiving a bespoke prompt containing the same content.

A minimal two-agent setup that works

For the pattern with the best effort-to-value ratio, build a generator and a critic. The generator gets the task, the tools, and a clear output spec. The critic gets the generator's output, the same spec, and one job: list every way the output fails the spec, with specifics. The generator then revises once against the critique. Two rounds maximum — more than that pays for diminishing returns. This setup catches a remarkable share of errors because generation and verification use different muscles, and it costs roughly twice a single run rather than five times. Run it on your highest-stakes recurring task first; if the critic consistently finds nothing, you do not need it there either.

Multi-agent workflows are a power tool: genuinely useful in the right hands, mostly a way to spend money in the wrong ones. The teams that succeed treat agents like contractors with written scopes, shared documentation, and a single accountable lead — not like a brainstorm that organizes itself. Start small, measure whether the second agent earns its keep, and expand only from demonstrated wins.

The orchestrator pattern, done properly

The most reliable multi-agent structure is the orchestrator with specialists: one agent plans and coordinates, others execute bounded subtasks. The orchestrator's job is decomposition and integration — it should never do the subtasks itself, because the moment it starts executing directly, the separation of concerns collapses and you have one confused agent with extra steps. Give the orchestrator three powers only: spawning workers with tight briefs, reading their results, and deciding what happens next. Workers get a self-contained brief — goal, constraints, what to return — and no knowledge of the larger plan beyond what they need. The integration step is where orchestrators earn their keep: results from parallel workers often conflict or overlap, and resolving that into a coherent output is genuine work. Keep worker briefs small enough that each worker's context stays fresh, and keep the orchestrator's context clean by having workers return summaries, not raw dumps.

  • The orchestrator never executes subtasks directly — it delegates, or the pattern is pointless.
  • Worker briefs are self-contained: goal, constraints, expected output format, nothing else.
  • Workers return summaries and conclusions, not full transcripts — the orchestrator's context is the scarcest resource.
  • The orchestrator validates integration: overlapping or conflicting worker results get reconciled explicitly, not concatenated.
  • Failure of one worker does not fail the run — the orchestrator retries, re-scopes, or proceeds with what is available.

Debugging when the swarm misbehaves

Multi-agent failures are harder to diagnose because the failure and its cause live in different agents. Build observability in from the start: every worker's brief, result, and the orchestrator's decisions should be logged where you can read them. When something goes wrong, diagnose in order. First, read the orchestrator's plan — was the decomposition sane, or did it split the task along the wrong lines? Most multi-agent failures are planning failures, not execution failures. Second, check the worker briefs — vague briefs produce vague results, and the orchestrator often under-specifies when it is itself confused. Third, read the failing worker's trace — did it misunderstand the brief, hit a tool problem, or run out of context? Fourth, check the integration — sometimes every worker did fine and the orchestrator mangled the merge. Fix at the level where the failure actually occurred; the temptation is to add more agents or more instructions, when the real fix is usually a sharper brief or a simpler decomposition.

Knowing when to collapse back to one agent

Multi-agent setups accumulate like sediment — added for one hard task, kept out of inertia. Regularly ask whether the complexity still earns its keep. Collapse back to a single agent when the subtasks turned out to be tightly coupled and workers constantly need each other's results, when one worker does ninety percent of the work and the orchestration is overhead on a single-agent task, when debugging takes longer than the task itself, or when costs exceed the value of the speedup. The collapse test is simple: run the task with one good agent and a clear prompt, and compare quality, time, and cost honestly. Single agents with good instructions and tools beat mediocre swarms consistently. Multi-agent is a tool for specific structural problems — parallelizable, decomposable work — not a default architecture.

A concrete cost picture

Put numbers on a typical comparison. A single agent handling a research task might take forty minutes and a few dollars of compute, working sequentially through sources. A five-worker swarm with an orchestrator might finish in twelve minutes — but the orchestrator's context carries five result summaries, each worker re-reads shared background, and coordination overhead means total compute is often three to five times the single-agent cost. Whether that trade is worth it depends on what you value: if a human is waiting and the answer is time-sensitive, the speedup justifies the cost. If the task runs overnight unattended, the single agent is cheaper and often more coherent. The mistake is assuming parallel is always better — it is faster in wall-clock time and almost always more expensive in total resources. Make the trade deliberately, per task, not as a standing architectural decision.

The critic pattern: agents reviewing agents

A useful addition to the orchestrator pattern is a dedicated critic: an agent whose only job is reviewing worker outputs before the orchestrator accepts them. The critic checks for factual errors, internal contradictions, missed requirements, and quality below the bar — the verification step that orchestrators often skip when they are busy integrating. This pattern earns its keep on high-stakes outputs where errors are expensive: customer-facing content, code that will be merged, analysis that drives decisions. It does not earn its keep on low-stakes drafts, where the critic's compute costs more than the occasional error it catches. Configure the critic with the same rigor as the workers: explicit review criteria, examples of past mistakes to watch for, and authority to reject work back to the worker with specific feedback. The most common failure is a critic that approves everything — counter it by sampling the critic's approvals for human audit, the same way you would audit any reviewer.

  • Give the critic explicit review criteria and examples of past mistakes, not vague instructions to check quality.
  • Grant rejection authority: the critic sends work back with specific feedback, not silent approval.
  • Audit the critic itself by sampling its approvals — a critic that approves everything is decoration.
  • Reserve the pattern for high-stakes outputs where error cost exceeds the critic's compute cost.
  • One critic per workflow is usually enough; critics reviewing critics is recursion, not rigor.

Keep reading