Harness pattern · Coordinate
Sequencing dependent tool calls
Make dependencies explicit data: chain steps with gates, run independent calls in parallel, and checkpoint between steps so a late failure costs one step.
The failure it prevents
In production: Calls that depend on each other race in parallel, or everything runs serially and the task takes ten times longer than it should.
Step two needs step one's output, but nothing wrote that down, so the harness fires both and step two guesses its input. The failure is intermittent, order-dependent, and impossible to reproduce on purpose. The opposite failure is slower and just as common: a task with twelve independent reads runs them one at a time because the first version was written sequentially and nobody revisited it.
When orchestration lives inside the model's head and the prompt, every run re-derives the plan. Sometimes it derives a worse one.
How it works
Name the dependency or lose it. If a call consumes another call's output, that is data: write it as an edge between steps. Anthropic's workflow patterns formalize the pieces: prompt chaining runs steps in sequence with programmatic gate checks between them, routing sends different inputs down different paths, and parallelization splits work that does not depend on itself.
Parallelize only true independence. In Anthropic's multi-agent research system, parallel tool calls (three or more at once from a subagent) cut research time substantially, and subagents run concurrently by design. The qualifier matters: those calls do not consume each other's output. The moment one does, it joins a chain.
Gate between steps. A chain gate checks the previous output (did the search return anything? did the write verify?) before spending the next call. LangGraph routes flow with the same primitives: state carries outputs forward, conditional edges branch on them, and Command objects send work to named next nodes, including after an error.
Checkpoint at step boundaries. Combined with durable execution, a failure at step nine restarts at step nine with steps one to eight intact, instead of re-running the whole plan against tools that may no longer answer the same way.
Mind parallel approvals. The OpenAI Agents SDK documents a subtlety worth copying: with guardrails running in parallel, the agent can keep working before a cancellation lands; a blocking mode exists for tools where that race is unacceptable.
When to use it
- Any task with more than three tool calls, where order has any consequence.
- Read-many workloads (fan-out research, bulk lookups) where parallel calls are free wins.
- Workflows mixing reads and writes, where the write must follow a verified read.
- Long chains where re-running from the top is expensive or changes answers.
When to skip or soften it
- Two-call tasks the model can order itself reliably; the graph machinery is heavier than the problem.
- Exploratory sessions where the next step genuinely depends on reading the last result with fresh eyes; keep the human or the model in the loop step by step.
Tradeoffs
Explicit graphs freeze a plan that the situation might want changed; a model routing freely adapts and a frozen edge does not. Keep the edges to the dependencies that are structural (data flow, safety order) and let the model choose inside steps. Parallel fan-out trades context too: five parallel search calls return five results into the same window, which is where context budgeting earns its keep.
A step graph with a gate (AgentsUse shape)
Dependencies as data: the read cannot start until the search lands and passes its gate. Independent steps with no shared edge may run in parallel.
# steps as data (shape; adapt to your runner)
steps:
- id: search
tool: web_search
- id: read_top_sources
tool: web_fetch
after: [search]
gate: "search.result_count >= 1"
- id: draft_summary
after: [read_top_sources]
gate: "read_top_sources.sources >= 1"Shape source: AgentsUse suggested shape; patterns per Anthropic, Building effective agents. Placeholders only; adapt names and numbers to your stack.
Implementation checklist
- Every multi-call task has its dependencies written down as steps and edges, not held in a prompt.
- A call runs in parallel only if it consumes no other call's output.
- Writes sit downstream of the reads that verify them, with a gate between.
- Each step checkpoints on completion so a late failure restarts at the failed step.
- Approval races are considered for parallel gated calls (blocking mode where needed).
- Fan-out width is capped, and returned results pass through the context budget.
Sources and freshness
- Anthropic: Building effective agents (chaining, routing, parallelization)
- Anthropic: How we built our multi-agent research system (parallel tool calls)
- LangGraph fault tolerance (error handlers and Command routing)
- OpenAI Agents SDK: guardrails (parallel versus blocking)
Claims on this page checked against these sources on 2026-10-08. Code and config blocks are shapes to adapt, not benchmarks.