In short

Adding agents to an AI system can degrade performance rather than improve it. A study by Google Research and DeepMind testing 180 configurations shows that multi-agent systems consume between 4 and 220 times more tokens than a single agent — often with no accuracy gain. The rule is not “more agents = better performance”, but “more agents = better performance on the right tasks, under the right conditions”.

Picture a kitchen: a chef alone prepares a simple dish in twenty minutes. Adding five assistants for the same dish multiplies coordination meetings, plate handoffs, and cross-checks — without improving the final taste. If the dish becomes a banquet with ten independent preparations, the team regains its advantage. Multi-agent systems obey the same principle: coordination has a fixed cost that only pays off when the task is genuinely parallelizable.


The baseline paradox

The “baseline paradox” refers to a well-documented phenomenon: a properly configured single agent often outperforms more complex multi-agent architectures on the same tasks.

Kim et al. (Google Research / DeepMind / MIT, 2025) evaluated 180 controlled configurations covering five different architectures — single agent, independent agents, centralized coordination, decentralized coordination, hybrid — tested on three model families (GPT, Gemini, Claude) and four benchmarks. Their results are clear:

  • Beyond a single-agent accuracy threshold of around 45%, multi-agent coordination produces diminishing or negative returns.
  • On sequential planning tasks (PlanCraft benchmark), all multi-agent variants degrade performance by 39 to 70% compared to the single agent.
  • Independent agents amplify errors 17.2 times, versus 4.4 times for centralized coordination.

The interpretation: when a single agent already performs well, adding agents mainly introduces complexity — additional error paths, costly coordination, and fault propagation between agents.

In plain terms : above 45% accuracy for the single agent, stacking agents adds almost nothing to the result but multiplies failure points. On tasks that must be done in order (sequential planning), multi-agent drops scores by 39 to 70%.


What the numbers say

Token overhead

Coordination cost is the first obstacle. Kim et al. quantify an overhead of 4 to 220 times more tokens for multi-agent systems depending on the task. This 220× factor is not an outlier: it reflects inter-agent communication, verification steps, and structural redundancy.

Xu et al. (2026, ICLR) put concrete figures on this gap for HumanEval: $0.020 per task for a single agent versus $0.026 for an equivalent multi-agent workflow — a 30% cost premium for near-identical performance (92.1% vs 91.6% pass@1).

BenchmarkSingle agentMulti-agent
HumanEval92.1%91.6%
MBPP81.4%81.1%
GSM8K93.3%93.0%
HotpotQA73.5 F173.5 F1

Xu et al.’s conclusion is direct: for workflows where all agents use the same base model (homogeneous workflows), a single agent in a multi-turn conversation delivers the same results at lower cost.

In plain terms : on standard tasks (code, math, QA), multi-agent costs 30% more to land within 0.5 points of the single-agent score. The premium pays for coordination, not for quality.

Identified failure modes

Cemri et al. (UC Berkeley / MIT, 2025) analyzed more than 1,600 execution traces across seven popular multi-agent frameworks and identified 14 distinct failure modes, grouped into three categories: system design flaws, agent misalignment, and task verification failures.

One pattern recurs often in “bag of agents” architectures — agents without a structured topology: an error from a first agent corrupts the context of the next, which produces an off-topic result passed to the third. The error propagates and amplifies at each step.

In plain terms : a multi-agent system can break through three distinct paths — bad topology design, agents contradicting each other, or missing final verification. Without a clear topology, an error in the first link contaminates the entire chain.


When multiple agents deliver real gains

The paradox has its limits. Several domains document clear gains.

Parallelizable tasks: Kim et al. measure +80.8% performance for centralized coordination on Finance-Agent, a financial processing benchmark where subtasks are independent of each other. Real parallelization justifies the overhead.

Pharmaceutical research: DrugAgent, a specialized multi-agent system, reaches 92% accuracy on complex drug discovery tasks, with a 4.92% ROC-AUC improvement over a single agent on the BindingDB benchmark for drug-target interaction prediction [UNVERIFIED: partial cross-source].

Long-horizon document research: multi-agent architectures (planning, parallel retrieval, synthesis) are documented as superior to simple LLM calls when the task requires simultaneously exploring many independent sources — as is the case with Deep Research tools (Google, Anthropic, OpenAI).

Specialized systems: DeepMind’s AlphaProof and AlphaGeometry 2 solve 4 out of 6 IMO 2024 problems. These specialized agent pipelines illustrate what works: agents with distinct roles, parallelizable tasks, real specialization.

In plain terms : multi-agent wins when the task naturally splits into independent subtasks, or when each agent plays a genuinely distinct role (real specialization, not cosmetic).

Single vs multi-agent decision matrix

ContextRecommendationWhy
Single task, <45% single-agent accuracyMulti-agent coordinationDocumented accuracy gains, cost justified (Kim et al., 2025).
Single task, ≥45% single-agent accuracyOptimized single-agentDiminishing multi-agent returns, 4-220× more tokens (Kim et al., 2025).
Sequential planning taskSingle-agentMulti-agent degrades 39-70% on PlanCraft (Kim et al., 2025).
Homogeneous workflow (same model everywhere)Multi-turn single-agentNear-identical performance, 30% premium for multi-agent (Xu et al., 2026).
Parallelizable task (independent subtasks)Multi-agent centralized coordination+80.8% on Finance-Agent, real parallelization (Kim et al., 2025).
Long-horizon document researchMulti-agent (plan → parallel → synthesis)Simultaneous exploration of independent sources, Deep Research pattern.
Costly errors, consensus requiredMulti-agent with votingDampens hallucinations through triangulation.
”Bag of agents” architecture without topologyTo avoid14 documented failure modes, amplified error propagation (Cemri et al., 2025).

The SciAgent lesson

In November 2025, the SciAgent paper (arXiv:2511.08151) claimed gold-medal performance at IMO 2025 through a hierarchical multi-agent system. It was retracted by its authors the same month, citing “the need for additional expert evaluation and major methodological adjustments”.

SciAgent illustrates a risk specific to the field: multi-agent performance claims are difficult to verify independently, the benchmarks used vary across studies, and the absence of a common standard encourages favorable comparisons. MultiAgentBench (ACL 2025) attempts to fill this gap, but has not yet been adopted as a reference.


Key takeaways

  • A well-configured single agent often outperforms multi-agent systems on sequential tasks — documented degradation of 39 to 70% (Kim et al., 2025).
  • The token overhead of multi-agent systems ranges from 4 to 220 times that of a single agent depending on the task.
  • The empirical threshold: beyond ~45% single-agent accuracy, multi-agent gains disappear.
  • Real gains are documented on parallelizable tasks, pharmaceutical research, and long-horizon document research.
  • Homogeneous workflows (same base model) can be simulated by a single agent with the same performance and lower cost (Xu et al., 2026).