What is multi-agent orchestration?
Multi-agent orchestration coordinates several agents toward one outcome, deciding who handles which step and how their separate results combine. The coordination itself is the straightforward part. The hard part is deciding what happens when one participant returns something plausible and quietly wrong.
The idea is intuitive because it mirrors how organizations already work. Rather than one generalist handling a complex case end to end, you use specialists: one agent that reads documents, one that checks a policy, one that drafts a response, and something deciding who does what and in what order.
It sits inside the wider coordination question covered at AI orchestration, which handles anything participating in an AI-driven process, including models, deterministic rules, system APIs, and human steps. Multi-agent orchestration is the narrower case where the participants are themselves agents, each choosing its own actions.
That narrowing is what creates the difficulty. Coordinating deterministic components is a solved problem in software. Coordinating components that each decide their own path, any of which can be confidently wrong, is not.
Single-agent systems fail visibly. Multi-agent systems fail plausibly. When one agent goes wrong, it usually produces something obviously broken or nothing at all. When one agent in a chain of five returns something plausible and incorrect, the remaining four accept it and build on it, and the system delivers a confident composite answer with a bad premise buried three steps back. Nothing errors. That difference is most of what this page is about.
Do you actually need more than one agent?
Less often than vendors suggest. Multiple agents earn their place when a task needs genuinely different expertise at different steps, when parts can run in parallel, or when context windows or tool permissions have to be kept separate. Otherwise one agent is cheaper and more reliable.
This section exists because multi-agent architectures are routinely presented as more advanced rather than as a trade. They are more capable and less reliable, and the trade is worth making deliberately rather than by default.
Six reasons people split, and which three survive scrutiny
| Reason | Verdict | The test that settles it | Cost of getting it wrong |
|---|---|---|---|
| Expertise | Holds up. Different kinds of reasoning at different steps | Would you give the two steps to different people? | One agent asked to do both does neither of them well |
| Parallelism | Holds up. Independent work that can run at once | Does either step actually depend on the other? | Sequential latency you had no reason to pay for |
| Separation | Holds up. Keeping context or permissions apart | Should any single agent hold every credential? | One agent with the whole permission set and the whole context |
| Complexity | Does not. Long is not the same as varied | Is the work genuinely varied, or just lengthy? | Handoffs introduced where none were needed |
| Org chart | Does not. Intuitive and misleading | Does the split follow the work or the reporting line? | Coordination overhead that mirrors your internal politics |
| Demos | Does not. Compelling and uninformative | What happens on the ten-thousandth run? | An impressive pilot that never reaches production |
The useful default is one agent, expanded only when a specific reason from the first list applies. Start with two before you attempt five, because the second agent is where you discover whether your handoffs, context passing, and failure handling actually work.
Why does reliability fall as agents multiply?
Because component reliability multiplies rather than averages. A chain where each agent succeeds ninety-five percent of the time does not succeed ninety-five percent of the time end to end. Five such agents produce roughly seventy-seven percent, and the decline accelerates.
The arithmetic below is the single most useful thing to internalise before designing a multi-agent system, because it runs against intuition. People reason about chains by averaging and the mathematics compounds.
| Agents | Each 95% reliable | Each 99% reliable | What it means in practice |
|---|---|---|---|
| 1 | 95% | 99% | One in twenty runs needs attention, or one in a hundred |
| 3 | 86% | 97% | The 95% chain is already failing one run in seven |
| 5 | 77% | 95% | Nearly a quarter of runs fail, from components that each look fine |
| 10 | 60% | 90% | Two chains, same length, wildly different outcomes |
Read the last row across and the design implication is clear. Moving each agent from ninety-five to ninety-nine percent takes a ten-agent chain from sixty percent to ninety percent. Per-component reliability matters far more than architecture, which is why effort spent hardening individual agents beats effort spent redesigning the coordination graph.
This is arithmetic, not measurement. Real chains include retries, recovery steps, and failures that are correlated rather than independent, so the true figure is neither this pessimistic nor as good as a demonstration implies. Treat the table as the shape of the problem rather than a forecast. The shape is what matters: reliability compounds downward, and it does so considerably faster than intuition suggests.
Four levers that recover the loss
- Verification steps. Insert a check between agents rather than trusting a handoff, particularly where one agent's output becomes another's premise. A cheap deterministic validation beats an expensive downstream failure.
- Bounded scope per agent. A narrowly scoped agent is more reliable than a broad one, so splitting for separation of concerns can raise per-component reliability rather than lowering it. This is the case where more agents helps.
- Human checkpoints on the irreversible. Place them by reversibility rather than by step count. One checkpoint before an action that cannot be undone recovers more reliability than three checkpoints on reads.
- Idempotent retries. If a step can be safely repeated, a transient failure costs a retry instead of a run. If it cannot, design that in before you need it, because retrofitting idempotency onto an agent that issues payments is a project.
What are the five coordination problems?
Task decomposition, agent selection, context passing, conflict resolution, and failure handling. Every multi-agent design solves all five of them, either explicitly or by accident, and the last two are the ones most often left entirely to chance rather than decided.
Naming them separately is useful because they fail differently and are solved in different places. Patterns for the coordination itself, such as supervisor and router topologies, are covered on the AI orchestration page.
- 01
Task decompositionBreaking the objective into steps that can be assigned. Fails when: the split is drawn along organizational lines rather than along genuine boundaries in the work, producing handoffs that carry more context than they can express.
- 02
Agent selectionDeciding which agent handles a given step. Fails when: two agents could plausibly handle it and the choice is made by ordering rather than by fit, which produces results that vary between runs for no visible reason.
- 03
Context passingDeciding what each agent receives at a handoff. Fails when: everything is passed, which reintroduces the noise the split was meant to avoid, or too little is passed, which means the receiving agent guesses. See context engineering.
- 04
Conflict resolutionDeciding what happens when two agents disagree. Fails when: nobody decided, so the orchestrator silently takes the first answer, the last answer, or the longest one. A disagreement between agents is information, and discarding it is a design choice usually made by omission.
- 05
Failure handlingDeciding what happens when one participant fails or returns something wrong. Fails when: the design assumes success, so a partial result propagates as though complete. This is the problem the next section covers, because it deserves more than a bullet.
What happens when one agent fails?
Four things can happen and only one of them is acceptable. The system can stop, retry, proceed without that contribution, or proceed as though it succeeded. The fourth is the default in most designs, and it is the one that produces confident wrong answers.
Partial failure is where multi-agent systems differ most from single-agent ones, and it is the least designed part of most implementations. The four possible responses are worth deciding explicitly, per step.
- Stop and escalate. The safest option and the right one where the missing contribution is load-bearing. Requires somewhere to escalate to and a person who will look.
- Retry. Correct for transient failures and dangerous for anything with a side effect. Only safe where the step is idempotent, which is a property you design in rather than discover.
- Proceed and mark the gap. Continue with the contribution flagged as missing, so downstream steps and the final output both know. Underused, and often the right answer.
- Proceed silently. The default when nobody decided. The output is delivered as complete when it is not, which is precisely the failure that looks like success.
A distinction worth holding onto: an agent returning an error is the easy case, because something can react to it. The hard case is an agent returning a confident answer that is wrong, because nothing downstream can tell the difference. Verification between agents addresses the second case and error handling does not, which is why the two are separate design activities rather than one.
The trace requirements follow directly. Reconstructing which participant introduced a bad premise needs per-agent, per-step records, and reconstructing it afterwards from a summary is generally impossible. See AI agent observability.