Multi-Agent Systems: Routing, Handoffs, Orchestration

Multi-agent systems compose familiar model and tool loops through explicit routing, delegation, ownership, and result contracts; extra agents help only when those boundaries provide a concrete benefit.

  • Explainer
  • 7 min read
Illustration of multiple agent roles coordinated through routing, handoffs, and orchestration boundaries.

A task grows broad enough to contain research, calculation, policy checks, and a final decision. It is tempting to label four agents, assign each a personality, and call the architecture advanced.

That changes the number of moving parts. It does not establish that the work has been divided well.

A multi-agent system is an application design that coordinates multiple agent roles or loops. Each role may receive different instructions, context, tools, permissions, and expected outputs. The roles can use the same foundation model. Agent identity is usually an application-level contract, not a distinct kind of model mind.

The engineering question is therefore not how many agents can we add? It is:

Can this work be separated while preserving the information, authority, and evidence each part needs?

Ownership patternsAgent count does not decide who owns the result
01

Single loop

Owns the replyAgent runtime
Bounded capabilityModel + tools

One runtime can plan, call tools, validate, and answer without another agent boundary.

02

Handoff

Routes the taskTriage agent
Ownership movesSpecialist agent

The specialist takes over the next response and the responsibilities attached to it.

03

Manager with specialist

Keeps reply ownershipManager agent
requestresult
Runs bounded taskSpecialist agent

The manager requests a result, then integrates or rejects it before answering.

Agent roles are application-level contracts. The boxes may use the same foundation model with different instructions, context, tools, permissions, and output requirements.

Start with the layers

The terms overlap in product documentation, but they should not collapse into one another.

  • A model processes represented input and produces output.
  • An agent runtime surrounds one or more model calls with tools, state, loops, permissions, checks, and stopping rules.
  • A worker or subagent is a bounded agent role invoked for a responsibility within a larger workflow.
  • An orchestrator is host-side control logic that may route work, schedule workers, pass state, validate results, enforce limits, choose recovery, and stop execution.
  • A workflow is the overall sequence and dependency structure through which work moves.

An orchestrator does not have to be another language model. It can be ordinary application code, a queue consumer, a state machine, a graph runtime, a model-guided supervisor, or a mixture of these.

Nor does every multi-agent design require one central orchestrator. Some systems use explicit graphs, peer handoffs, queues, or event-driven coordination. The labels in this field are conventions rather than universal standards. What matters is where control and responsibility actually sit.

One capable loop is the baseline

A single agent can already use multiple tools, retrieve specialized context, run validation, and ask a human for approval. Keeping that work inside one loop often preserves context and makes failures easier to inspect.

A new agent boundary is justified only when it buys something concrete, such as:

  • a responsibility with a genuinely different input or output contract;
  • a narrower context that improves focus or protects unrelated data;
  • distinct tools or permissions;
  • work that can safely run in parallel;
  • an explicit isolation or approval boundary;
  • independent execution whose artifacts can be checked separately.

Roles such as planner, research worker, execution worker, and validator describe engineering responsibilities. Elaborate personas are not required for specialization. Different instructions, data access, tools, permissions, and acceptance criteria do the substantive work.

Ownership separates common patterns

Several designs can all be called “multi-agent” while behaving differently.

Routing selects a path

A router classifies a request and chooses a specialist, workflow, model, or capability path. A support system might send account-security questions to a constrained workflow and product questions to a retrieval workflow.

Routing does not necessarily add multiple active agents. It may choose exactly one path. The production obligations are the route criteria, uncertain-route behavior, permissions on each path, and fallback when the selected dependency fails.

A handoff moves ownership

In a handoff, one agent transfers the next response or task to a specialist. The specialist receives a defined context and becomes responsible for the next result.

That can be useful when the specialist needs a different policy or conversation responsibility. It also means the original coordinator may no longer control every decision inside that branch. The handoff must say what context moves, what authority moves, and how completion or failure returns.

A manager can keep ownership

In a manager-specialist pattern, the manager invokes a specialist as a bounded capability. The specialist might return research findings, a calculation, or a draft. The manager remains responsible for integrating, validating, and presenting the final result.

This is sometimes called “agents as tools.” The phrase is implementation-specific, but the ownership distinction is durable: assistance returns to the manager rather than transferring the user-facing reply.

Orchestrator-worker designs decompose work

An orchestrator may create subtasks, assign them to workers, collect results, resolve dependencies, and decide what happens next. This can help when work has separable outputs, but it creates interfaces that a single loop did not need.

The orchestrator still needs a stopping rule. “The model will know when it is done” is not a production contract. Surrounding code may enforce terminal states, task-completion conditions, iteration limits, time or cost budgets, failure thresholds, or human approval.

Decomposition and parallelism are different decisions

Breaking work into parts does not make those parts independent.

Sequential:  research -> analyze -> publish

Parallel:    source A --\
             source B ----> join -> compare
             source C --/

Conditional: classify -> billing path
                      \-> security path

The first flow can use three agents and remain sequential. The second can use three concurrent calls, but only if each source can be inspected without results from the others. The third chooses one branch.

Before parallelizing, identify hidden dependencies. Does each worker have enough context? Can its result be interpreted independently? Can the join detect overlap, contradiction, and missing coverage? If not, fan-out may produce faster fragments that cannot be assembled reliably.

Parallel work can also increase peak rate limits and cost even when it reduces wall-clock latency. Agent count and workflow topology are separate architecture choices.

Agent boundaries are interfaces

Natural-language messages are flexible, but flexibility alone does not make coordination reliable. When one component consumes another’s work, the return channel benefits from an explicit contract.

Depending on the task, a result may carry:

Field Why it matters
Task identifier and status Connect the result to the requested work and make incomplete states visible
Findings or output Provide the bounded result the caller needs
Evidence and artifacts Let the caller inspect sources, files, tests, or tool results
Limitations State missing inputs, unresolved conflicts, or confidence constraints
Errors Separate a failed step from a successful empty result
Required next action Make retry, escalation, approval, or integration explicit

This is not a universal schema. A calculation tool and a research worker need different contracts. The principle is to make downstream assumptions inspectable.

A response can match a JSON Schema and still be wrong. Format validity proves that the interface shape is acceptable; it does not prove semantic correctness. Evidence, deterministic checks, tests, authoritative sources, or human review may still be required.

Coordination has a bill

Every boundary can add:

  • another model call and another context assembly;
  • tokens, latency, and provider load;
  • information lost or distorted in transfer;
  • duplicated or contradictory work;
  • state that must be stored and reconciled;
  • permissions and credentials that must be scoped;
  • another place to retry, time out, or stop;
  • more traces to interpret during an incident.

Errors can propagate while the final output remains fluent:

Worker A returns a wrong premise
-> Worker B accepts it as input
-> Manager integrates B's coherent result
-> the final answer looks consistent but rests on A's error

Adding a reviewer agent does not automatically repair this chain. Two agents may use the same model, source, instructions, or false assumption. Review can catch some mistakes; independent verification uses stronger evidence such as tests, schemas, source checks, deterministic tools, or appropriately independent human judgment.

Choose the smallest architecture that preserves the contract

Consider a product that answers billing, technical, and account-security questions.

A single agent with constrained tools may be enough when the policies are compact and the same runtime can enforce every permission. Routing may help when security requests require a distinct approval path. A handoff may fit when a specialist should own the continued conversation. A manager-specialist pattern may fit when one assistant should synthesize bounded findings into a consistent response.

None is universally more advanced. The useful design makes these questions answerable:

  1. Who owns the next decision and the final result?
  2. What context and state does each participant actually receive?
  3. Which tools and permissions can each participant use?
  4. Is the work sequential, parallel, or conditional, and why?
  5. What must cross each handoff, and how is it validated?
  6. What happens when a route is uncertain or a worker fails?
  7. What event stops the workflow?
  8. Which observable evidence lets an operator reconstruct what happened?

If the answers stay vague, more agents have added architecture without adding a dependable boundary. Multi-agent design becomes useful when decomposition clarifies ownership and control enough to justify its coordination cost.

References

  1. Orchestration and handoffsOpenAI
  2. Building effective agentsAnthropic