In brief

In 2026, Claude Code offers four distinct mechanisms for coordinating multiple LLM instances on a single task. Each has a different coordination model, different guarantees, and different limitations. Understanding these distinctions is the difference between a pipeline that holds in production and one that fails silently.

In short: multi-agent orchestration in 2026 is like choosing between four team-building formations: a wave of independent interns delivering in parallel (Agent Tool), a team with a shared task list (Agent Teams), specialists in separate offices (Session Agents), or observers monitoring without intervening (MCP Hooks). Each solves a different problem — confusing them means picking the wrong structure.


The four mechanisms

Agent Tool — parallel dispatch

The Agent Tool (also called Task Tool in the documentation) is the base mechanism for launching independent sub-agents. The parent formulates N atomic tasks, launches N instances in parallel, waits for their results, and consolidates.

Each sub-agent runs in its own context, with its own tools and permissions. The official documentation is explicit: “Each subagent invocation creates a new instance with fresh context.” There is no shared token pool — field measurements show that two parallel agents consume 135K tokens combined without mutual compaction.

Current parameters (March 2026)

The model parameter accepts aliases sonnet, opus, haiku, inherit, or a full identifier (claude-opus-4-6). The context ceiling per sub-agent has risen to ~128K tokens with Opus 4.6 and Sonnet 4.6 (v2.1.77). The output limit has been multiplied by 8: 64K tokens instead of 8,192 for Opus 4.6.

Active constraint to watch: an unresolved bug (#29231) causes the [1m] suffix to be dropped when spawning sub-agents — a parent running on 1M tokens spawns children limited to the standard context (~200K). The bug is open with no Anthropic response to date.

Actionable failure signal: when a sub-agent fails, the parent receives a parseable text response. This is the most useful property of this mechanism — unlike Agent Teams, the failure is visible.

Structural constraint: no nesting. A sub-agent cannot spawn its own sub-agents. The hierarchy is flat: one level of depth maximum.

Rate limits: 12 parallel Opus agents consume approximately 168K tokens per minute. On Tier 1 (30K ITPM), that is 5.6 times the limit. Massive parallelism requires Tier 2 at minimum.

In short: the per-sub-agent token budget is generous (≈ 100K Sonnet, ≈ 128K Opus), but the rate limit is shared across all simultaneous agents. With 12 Opus running in parallel, you exceed 5× the Tier 1 limit — either you are on Tier 2+, or you stagger launches over time. Otherwise, agents fail in cascade.


Agent Teams — coordination with shared state

Agent Teams is a different mechanism: a “lead” agent coordinates “teammates” who share a common task list and can send each other messages.

What distinguishes Agent Teams from Agent Tool:

  • Shared task list: the lead deposits tasks, teammates claim them (including self-assignment after completing a previous task).
  • File locking: prevents race conditions when two teammates attempt to claim the same task simultaneously.
  • Bidirectional mailbox: via SendMessage, lead and teammates can exchange messages. Since v2.1.77, a SendMessage to a stopped background agent automatically restarts it.
  • Plan approval mode: a teammate can be constrained to read-only until approved by the lead.

Active bugs limiting its use

Three serious bugs are open on the GitHub repository as of March 2026:

BugImpact
#32368 — broken model inheritanceTeammates use hardcoded Opus, not the lead’s model. The v2.1.72 fix is not functional for custom endpoints and Bedrock.
#27555 — indistinguishable messagesMessages sent via SendMessage display with the prefix ⏺ Human:, indistinguishable from user input. Has caused documented destructive actions (#28627).
#24316 — no specialized agentsImpossible to use agents with defined profiles (restricted tools, model specific to role) as teammates. All teammates are identical generic instances. 27 comments, active March 18, 2026.

Ephemeral by design: tasks are purged after completion. /resume does not restore teammates — after a resume, the lead may attempt to contact agents that no longer exist.

Cost: the official documentation estimates that Agent Teams consumes approximately 7x more tokens than a standard plan-mode session, because each teammate maintains its own context. Broadcasting (sending a message to all teammates) costs N×M context injections — not recommended at scale.


Session Agents — isolation via worktree

The third approach consists of launching Claude Code instances in separate git worktrees. Each instance operates on an isolated branch, with its own files, without risk of collision with the others.

This mechanism is relevant when isolation is structural: multiple agents working on different parts of a codebase, with human review before merge. This is the pattern Anthropic used for its C compilation project — 16 parallel agents on ~100K lines, without mutual interference.

Documented constraint: this mode works reliably only with Opus. Field reports show that Sonnet and Haiku do not maintain behavioral consistency over long sessions in a worktree isolation context.


MCP Hooks — orchestration observability

Hooks are not a coordination mechanism per se, but they make orchestration observable and controllable. In March 2026, Claude Code exposes 22 hook events.

Two events are particularly relevant for orchestration:

  • SubagentStart: fires immediately after a sub-agent is spawned via Agent Tool. Asynchronous. Allows injecting additional context via additionalContext.
  • SubagentStop: fires after sub-agent completion, before the result is returned to the parent. Synchronous and blocking — can prevent stoppage and force continuation. Rich payload: transcript path, last message.

These two hooks were documented but unconfirmed for Agent Tool until early 2026. They are now explicitly documented as firing for sub-agents spawned via Agent Tool.

Other hooks relevant to orchestration: StopFailure (with matcher on error type: rate limit, authentication, billing, etc.), PreCompact/PostCompact, TeammateIdle, TaskCompleted.


What the literature says about reliability

Architecture outweighs the model

The MAESTRO research (arXiv 2601.00481, January 2026) evaluated 12 multi-agent systems with framework-agnostic traces. Its conclusion: system architecture is the dominant factor for reproducibility, cost, and accuracy — ahead of model choice. Run-to-run variance is substantial even with identical systems and tasks, and this variance is explained by structural design choices.

Princeton University and the Berkeley Klein Center published in February 2026 (“Towards a Science of AI Agent Reliability”, arXiv 2602.16666) an evaluation of 14 models on 12 reliability metrics: consistency, robustness, predictability, safety. Result: recent capability gains have produced only modest reliability improvements. Moving to the next model up does not fix structural reliability problems.

The root cause of failures: the probabilistic/deterministic mismatch

“Characterizing Faults in Agentic AI” (arXiv 2603.06847, March 2026) analyzed 13,602 issues and pull requests across 40 agentic system repositories. Its central conclusion: the majority of failures stem from a mismatch between probabilistically generated outputs and deterministic interface constraints.

The model generates probabilistic text — it never produces exactly the same output twice. The rest of the system expects deterministic inputs: valid JSON, an exact filename, a response in a precise format. When the interface between the two is poorly designed, the system breaks. Explicit format validations between agents — not a more powerful model — is what fixes this problem. 83.8% of practitioners surveyed confirm that this taxonomy covers their real-world cases.

Silent failures

“Silent Failures in Multi-Agent Systems” (arXiv 2511.04032, November 2025) formalizes a particularly difficult failure mode to detect: agents that drift, loop, or omit critical details — without generating an explicit error. These silent failures are detectable by machine learning (XGBoost at 98% precision on an annotated corpus), but not by simple inspection of the final result.

Measured failure rates

MASFT (UC Berkeley/CMU, arXiv 2503.13657) remains the reference for failure rates in real-world multi-agent systems: between 41% and 86.7% depending on task type and system. No recent paper contradicts these figures. Princeton/Berkeley (2602.16666) confirms them indirectly by showing that capability gains do not reduce them significantly.


What works, what doesn’t yet

Validated

  • 12 Sonnet agents in parallel without quality degradation or rate limits, on a Tier 2 account. Empirical tests show a total time of approximately 16 seconds for a parallel document research task.
  • Dispatch by cognitive type: choosing the model based on the nature of the operation (mechanical extraction to Haiku, nuanced analysis to Sonnet) is more effective than dispatching on the perceived “size” or “complexity” of the task. Three independent sources converge on this point (MAESTRO, Princeton/Berkeley, field observations).
  • Output to file: writing to a file rather than returning by text bypasses the output limit and allows post-execution verification. With Opus 4.6, the limit has risen to 64K tokens — the pattern remains recommended for outputs exceeding that bound.
  • SubagentStart/Stop hooks: confirmed as exploitable for observability and dynamic context injection.

In progress or unstable

  • Agent Teams: the mechanism exists and the primitives work (task list, mailbox, self-claim), but three active bugs (#32368, #27555, #24316) reduce reliability. Use with caution in production, and only with standard endpoints.
  • Specialized agents as teammates: the feature allowing use of agents with defined profiles (restricted tools, model-specific by role) is not yet implemented (issue #24316 open, 27 active comments). In the meantime, all teammates are generic instances.
  • 1M context for sub-agents: the parent can use 1M tokens, but sub-agents are limited to standard context (~200K) due to bug #29231. Unresolved.

Which mechanism for which need?

Usage contextRecommendationWhy
N independent parallel tasks, no shared state neededAgent ToolReliable, actionable failure signal, isolated budgets. The default mechanism
Dynamic decomposition with collaboration (a lead distributes, teammates self-assign)Agent Teams (with caution)Enables coordination, but 3 bugs were open in March 2026 — to be avoided in critical production
Multi-file refactoring without collisionSession Agents (worktrees)Physical git isolation, ideal for 16+ agents over ~100K lines — Opus only
Observability and control outside the flow (logs, context injection, gating)MCP HooksNo coordination, but SubagentStart/Stop make it possible to instrument without modifying the agents

Key takeaways

  • Four distinct mechanisms: Agent Tool (stateless parallelism), Agent Teams (coordination with state), Session Agents (worktree isolation), MCP Hooks (observability). Each mechanism has a different domain of application.
  • Agent Tool is the most reliable mechanism: actionable failure signal, independent budgets, parallelism up to 12 agents (Tier 2). Agent Teams adds coordination but with active bugs as of March 2026.
  • The literature converges: three independent teams (MAESTRO, Princeton/Berkeley, Berkeley/CMU) show that capability gains do not produce proportional reliability gains. Architecture comes first.
  • Failures come from interfaces: format mismatches between agents, silent failures, non-compliance with instructions without explicit error. Output validations between agents are more effective than a more powerful model.
  • Tier 2 is the minimum for parallelism beyond 5-6 Opus agents.