In short

A solo agent on Claude Opus 4.5 costs around $9 for a 20-minute session. The same work handed to a planner + generator + evaluator harness costs $200 over six hours — a 22× factor, measured by Anthropic on a concrete software generation case. Agents generally consume 4× more tokens than chat interactions; multi-agent systems consume 15× more. This added cost is not evenly distributed: it comes from a handful of specific, identifiable, and partially controllable mechanisms. And on standard software development tasks, empirical studies show that solo agents can outperform multi-agent architectures in quality — while costing three times less.

In short: think of a construction site. A single all-rounder craftsman who does everything is 9 cost units. Three specialized craftsmen passing plans, reviewing each other and correcting each other — that is 200 units, and the result is not necessarily better. Coordination is expensive, the context passed between craftsmen is expensive, and every check adds hours. Multi-agent only pays off when the solo agent really cannot do it alone.


The $9 → $200 case: what happened?

Anthropic published a detailed analysis in March 2026 of the “harness” architecture they use for long software development tasks. The starting point: a solo agent (Claude Opus 4.5) generates a retro game maker in a 20-minute session for around $9. The same category of task, handed to a full harness with a planner agent, a generator agent, and a QA evaluator agent, costs $200 over six hours.

A second documented run on a different task (an audio workstation) shows the internal breakdown (Anthropic Engineering, March 2026):

  • Planner: $0.46 — less than 0.4% of the total
  • Build phases (3 rounds of generation): $113.85 — 91.3% of the total
  • QA phases: $10.39 — 8.3% of the total
  • Total: $124.70 over 3h50

The surprise: orchestration itself is nearly free. The planner costs less than a dollar. What is expensive is the actual work — repeated code generation by the generator, call after call, with a context that grows larger each turn. Understanding the added cost therefore requires looking beyond “orchestration overhead” in the narrow sense.


Anatomy of costs: four distinct sources

1. Recreating the system prompt for each sub-agent

Each sub-agent is an independent API session. It instantiates its own system prompt from scratch, which means system instructions are billed as input tokens on every new call. For an 8,000-token system prompt on Claude Sonnet 4.6 ($3 per million input tokens), each agent adds $0.024 in fixed startup cost. Over fifty calls per agent in a long session, that amounts to $3.60 in pure recreation cost — for a single agent (Anthropic — Prompt caching documentation).

Prompt caching radically changes this calculation: the first write costs 1.25× the standard rate, but each subsequent read costs only 0.1× — an 85% reduction from the second call onwards. The breakeven is reached at the second call. Beyond three sub-agents sharing the same instruction prefix, the reduction on this component exceeds 70%.

The technical constraint: in parallel sub-agent deployment, the first agent must be launched with a slight delay (0.5 seconds is enough) to give it time to create the cache entry before the others query it.

2. Coordination overhead: supra-linear with the number of agents

A December 2024 study (arxiv:2512.08296) measured coordination overhead by architecture type, comparing token consumption against a solo agent:

ArchitectureOverhead vs. solo agentEfficiency (successes / 1,000 tokens)
Solo agent0%67.7
Independent multi-agents+58%42.4
Decentralized multi-agents+263%23.9
Centralized multi-agents+285%21.5
Hybrid multi-agents+515%13.6

The scaling law observed for interaction turns is T = 2.72 × (n + 0.5)^1.724, with R² = 0.974. In other words: double the number of agents, and interaction turns do not increase by 2× but by more than 3×. This supra-linear relationship is the structural explanation for the 15× factor documented by Anthropic for multi-agent systems compared to standard chat interactions.

3. Framework scaffolding: an invisible cost

The choice of orchestration framework is not neutral. A benchmark on a three-agent pipeline, processed at 10,000 requests per day with GPT-4o, measures a significant gap between LangGraph (~800 tokens per request, ~$32 per day) and CrewAI (~1,250 tokens per request, ~$50 per day). The 450-token difference corresponds to the automatic role definitions that CrewAI injects for each agent — around 150 tokens per agent of scaffolding that the user cannot control. [This measurement comes from a single source and should be treated as indicative.]

LangGraph offers granular control over prompts; CrewAI abstracts that layer at the price of systematic overhead. The choice of architecture thus encodes a recurring cost, invisible line by line but significant at scale.

4. Verification loops: the exponential cost of errors

An empirical benchmark across seven frameworks (arxiv:2511.00872) finds that the multi-agent system AgentOrchestra devotes 41.54% of its trajectory to correction attempts — versus an average of 4.81% for solo agents. Total cost ranges from $2.13 (solo agent GPTswarm) to $370.19 (AgentOrchestra multi-agents) for the same 1,200 software development tasks.

The underlying mechanism is structural. A Reflexion loop over ten cycles can consume 50× more tokens than a single linear pass, because each turn includes the full accumulated conversation history (Stevens Online). Turn 1 = 200 tokens. Turn 10 = 200 tokens plus nine times the size of each previous output. Context complexity is O(n²) for a solo LLM that self-corrects.

In short: the four cost sources have very different profiles. The system prompt sent to each sub-agent is free with prompt caching from the 2nd call (-85%). Coordination overhead, in contrast, scales supra-linearly — doubling the number of agents triples the interaction turns. Framework scaffolding is invisible but permanent. Verification loops are the most dangerous: they can multiply the bill by 50× without anyone noticing.


The counter-evidence: solo agents sometimes outperform multi-agents

A finding that appears across several independent studies: solo agents with optimized prompts can match or exceed multi-agent architectures on standard tasks, at a fraction of the cost.

The arxiv:2511.00872 study explicitly concludes that “single-agent systems consistently demonstrate better performance than multi-agent systems” on standard software development tasks. OpenHands, a solo agent, outperforms AgentOrchestra in output quality while costing $115 versus $370 for 1,200 identical tasks. The performance difference stems largely from the fact that a solo agent can maintain a coherent context across the entire task, whereas a multi-agent system must summarize and pass that context between agents — with the information loss that entails.

A field result points in the same direction: reducing the number of tools available to an agent by 80% improves its results (Vercel, cited by Aakash Gupta). Fewer options, less hesitation, fewer deliberation tokens over which tool to call. Simplifying each agent’s interface is often more effective than adding a coordination layer.

An ICLR 2025 workshop on foundation models in real-world conditions notes the same trend: “single agents with optimized prompts can match multi-agent performance while saving tokens.”

This counter-evidence does not mean that multi-agents is useless. It delimits the perimeter where it genuinely is useful — and serves as a reminder that architectural complexity carries its own cost that must be justified, not assumed.


When the added cost is justified

Anthropic states the decision criterion directly in its harness guide (Anthropic Engineering, March 2026): “The evaluator is worth the cost when the task exceeds what the model does reliably solo.” And, more fundamentally, in “Building Effective Agents” (2024): “Agentic systems often trade latency and cost for better performance — you need to ask when that trade-off makes sense.”

Three cases where the multiplier is empirically justified:

Real parallelization. Anthropic reports a 90.2% performance gain for its multi-agent research system (a lead Opus 4 coordinating Sonnet 4 sub-agents) compared to a solo Opus 4 agent, on tasks where searches can be conducted in parallel. Processing time is reduced, and coverage of the search space increases. This is the use case where multi-agents has the clearest advantage.

Exceeding the context window. When a task does not fit within a single context window, multi-agents is not one option among others — it is the only option. Google Research’s Chain of Agents architecture reduces complexity from O(n²) to O(nk) by splitting context across specialized agents, which justifies it economically for very long-context tasks.

Specialization with model cascading. The BudgetMLAgent project (IIT Bombay, AIMLSystems 2024) achieves a 94.2% cost reduction — from $0.931 to $0.054 per run — by routing mechanical sub-tasks to cheaper or free models, while achieving a higher success rate than solo GPT-4 (32.95% vs. 22.72%). Multi-agents can cost less than a high-end solo agent when dispatch is intelligent.

The depreciation problem

A practical warning: every component in a harness encodes an assumption about what the model cannot do alone. Anthropic states this directly in its March 2026 article — “Every component in a harness encodes an assumption about what the model can’t do.” On Opus 4.5, an external QA evaluator was necessary for certain tasks. On Opus 4.6, those same tasks fall within the model’s native competence window — the evaluator becomes unnecessary overhead.

An investment of “thousands of engineering hours” (Aakash Gupta) in a complex harness therefore depreciates with each new model generation. The pace of progress in base LLM capabilities is currently fast enough that some multi-agent architectures become obsolete before they have recouped their development cost.

The drop in LLM API prices — roughly 80% over twelve months according to pricepertoken.com — also changes the calculation: systems previously unviable at $200 per run may become acceptable at $40 with new pricing. But if this drop leads teams to proportionally increase the complexity of their architectures, the savings disappear. The actual ROI of complex harnesses is not quantified in the available literature, which is itself a warning signal for teams making heavy investments in these structures.


Key takeaways

  • A multi-agent system consumes up to 15× more tokens than a standard chat interaction; a full harness can cost 22× more than a solo agent on the same task (Anthropic Engineering, March 2026).
  • The main cost is carried by the worker agent (91% in the documented case), not by orchestration. The planner costs less than a dollar.
  • Prompt caching reduces context recreation cost by 85% from the second call onwards — it is the most immediate optimization lever for a harness in production.
  • On standard software development tasks, solo agents with optimized prompts generally outperform multi-agent architectures in both quality and cost (arxiv:2511.00872).
  • Multi-agents is justified in three specific cases: real parallelization of sub-tasks, exceeding the context window, and model cascading with intelligent dispatch.