In short
Adding agents to an AI system can degrade performance rather than improve it. A study by Google Research and DeepMind testing 180 configurations shows that multi-agent systems consume between 4 and 220 times more tokens than a single agent — often with no accuracy gain. The rule is not “more agents = better performance”, but “more agents = better performance on the right tasks, under the right conditions”.
Picture a kitchen: a chef alone prepares a simple dish in twenty minutes. Adding five assistants for the same dish multiplies coordination meetings, plate handoffs, and cross-checks — without improving the final taste. If the dish becomes a banquet with ten independent preparations, the team regains its advantage. Multi-agent systems obey the same principle: coordination has a fixed cost that only pays off when the task is genuinely parallelizable.
The baseline paradox
The “baseline paradox” refers to a well-documented phenomenon: a properly configured single agent often outperforms more complex multi-agent architectures on the same tasks.
Kim et al. (Google Research / DeepMind / MIT, 2025) evaluated 180 controlled configurations covering five different architectures — single agent, independent agents, centralized coordination, decentralized coordination, hybrid — tested on three model families (GPT, Gemini, Claude) and four benchmarks. Their results are clear:
- Beyond a single-agent accuracy threshold of around 45%, multi-agent coordination produces diminishing or negative returns.
- On sequential planning tasks (PlanCraft benchmark), all multi-agent variants degrade performance by 39 to 70% compared to the single agent.
- Independent agents amplify errors 17.2 times, versus 4.4 times for centralized coordination.
The interpretation: when a single agent already performs well, adding agents mainly introduces complexity — additional error paths, costly coordination, and fault propagation between agents.
In plain terms : above 45% accuracy for the single agent, stacking agents adds almost nothing to the result but multiplies failure points. On tasks that must be done in order (sequential planning), multi-agent drops scores by 39 to 70%.
What the numbers say
Token overhead
Coordination cost is the first obstacle. Kim et al. quantify an overhead of 4 to 220 times more tokens for multi-agent systems depending on the task. This 220× factor is not an outlier: it reflects inter-agent communication, verification steps, and structural redundancy.
Xu et al. (2026, ICLR) put concrete figures on this gap for HumanEval: $0.020 per task for a single agent versus $0.026 for an equivalent multi-agent workflow — a 30% cost premium for near-identical performance (92.1% vs 91.6% pass@1).
| Benchmark | Single agent | Multi-agent |
|---|---|---|
| HumanEval | 92.1% | 91.6% |
| MBPP | 81.4% | 81.1% |
| GSM8K | 93.3% | 93.0% |
| HotpotQA | 73.5 F1 | 73.5 F1 |
Xu et al.’s conclusion is direct: for workflows where all agents use the same base model (homogeneous workflows), a single agent in a multi-turn conversation delivers the same results at lower cost.
In plain terms : on standard tasks (code, math, QA), multi-agent costs 30% more to land within 0.5 points of the single-agent score. The premium pays for coordination, not for quality.
Identified failure modes
Cemri et al. (UC Berkeley / MIT, 2025) analyzed more than 1,600 execution traces across seven popular multi-agent frameworks and identified 14 distinct failure modes, grouped into three categories: system design flaws, agent misalignment, and task verification failures.
One pattern recurs often in “bag of agents” architectures — agents without a structured topology: an error from a first agent corrupts the context of the next, which produces an off-topic result passed to the third. The error propagates and amplifies at each step.
In plain terms : a multi-agent system can break through three distinct paths — bad topology design, agents contradicting each other, or missing final verification. Without a clear topology, an error in the first link contaminates the entire chain.
When multiple agents deliver real gains
The paradox has its limits. Several domains document clear gains.
Parallelizable tasks: Kim et al. measure +80.8% performance for centralized coordination on Finance-Agent, a financial processing benchmark where subtasks are independent of each other. Real parallelization justifies the overhead.
Pharmaceutical research: DrugAgent, a specialized multi-agent system, reaches 92% accuracy on complex drug discovery tasks, with a 4.92% ROC-AUC improvement over a single agent on the BindingDB benchmark for drug-target interaction prediction [UNVERIFIED: partial cross-source].
Long-horizon document research: multi-agent architectures (planning, parallel retrieval, synthesis) are documented as superior to simple LLM calls when the task requires simultaneously exploring many independent sources — as is the case with Deep Research tools (Google, Anthropic, OpenAI).
Specialized systems: DeepMind’s AlphaProof and AlphaGeometry 2 solve 4 out of 6 IMO 2024 problems. These specialized agent pipelines illustrate what works: agents with distinct roles, parallelizable tasks, real specialization.
In plain terms : multi-agent wins when the task naturally splits into independent subtasks, or when each agent plays a genuinely distinct role (real specialization, not cosmetic).
Single vs multi-agent decision matrix
| Context | Recommendation | Why |
|---|---|---|
| Single task, <45% single-agent accuracy | Multi-agent coordination | Documented accuracy gains, cost justified (Kim et al., 2025). |
| Single task, ≥45% single-agent accuracy | Optimized single-agent | Diminishing multi-agent returns, 4-220× more tokens (Kim et al., 2025). |
| Sequential planning task | Single-agent | Multi-agent degrades 39-70% on PlanCraft (Kim et al., 2025). |
| Homogeneous workflow (same model everywhere) | Multi-turn single-agent | Near-identical performance, 30% premium for multi-agent (Xu et al., 2026). |
| Parallelizable task (independent subtasks) | Multi-agent centralized coordination | +80.8% on Finance-Agent, real parallelization (Kim et al., 2025). |
| Long-horizon document research | Multi-agent (plan → parallel → synthesis) | Simultaneous exploration of independent sources, Deep Research pattern. |
| Costly errors, consensus required | Multi-agent with voting | Dampens hallucinations through triangulation. |
| ”Bag of agents” architecture without topology | To avoid | 14 documented failure modes, amplified error propagation (Cemri et al., 2025). |
The SciAgent lesson
In November 2025, the SciAgent paper (arXiv:2511.08151) claimed gold-medal performance at IMO 2025 through a hierarchical multi-agent system. It was retracted by its authors the same month, citing “the need for additional expert evaluation and major methodological adjustments”.
SciAgent illustrates a risk specific to the field: multi-agent performance claims are difficult to verify independently, the benchmarks used vary across studies, and the absence of a common standard encourages favorable comparisons. MultiAgentBench (ACL 2025) attempts to fill this gap, but has not yet been adopted as a reference.
Key takeaways
- A well-configured single agent often outperforms multi-agent systems on sequential tasks — documented degradation of 39 to 70% (Kim et al., 2025).
- The token overhead of multi-agent systems ranges from 4 to 220 times that of a single agent depending on the task.
- The empirical threshold: beyond ~45% single-agent accuracy, multi-agent gains disappear.
- Real gains are documented on parallelizable tasks, pharmaceutical research, and long-horizon document research.
- Homogeneous workflows (same base model) can be simulated by a single agent with the same performance and lower cost (Xu et al., 2026).