In brief

Chain-of-Debate (CoD) has multiple LLM instances debate the same question: each model proposes an answer, critiques the others’ responses, and revises its own over several rounds. Graph-of-Debate (GoD) extends this mechanism by structuring arguments as a non-linear graph rather than a sequential chain. The central idea: a bias or error present in one individual model is detected and corrected by the other debate participants. MASS (Multi-Agent System Search) goes further by automatically optimizing these debate architectures.

In short: a single lawyer drafting an opinion sees what their training has taught them to see. A bench of several judges arguing adversarially spots what a single one would have missed. Chain-of-Debate does the same with LLMs: you put several around a question, they argue, they attack each other’s arguments, and you take the decision that survived the objections. The cost: N models × R inference rounds — that many times the bill of a simple answer.


From individual reasoning to collective reasoning

Chain-of-Thought (CoT), introduced as a prompting technique, forces a single LLM to decompose a problem into explicit intermediate steps. This approach is individual: one model reasons, and its biases are its own.

CoD starts from a different premise: if multiple independent models reason about the same question, their errors will not be correlated. Divergences between agents signal zones of uncertainty. The emerging consensus is more robust than an individual answer, even one from a larger model.

The basic structure of CoD operates in three steps:

  1. Initial proposal: each agent produces its answer and reasoning independently.
  2. Cross-critique: each agent reads the others’ responses and formulates its objections.
  3. Revision: each agent updates its position taking into account the critiques received.

These rounds repeat until convergence or until a predefined maximum number of rounds.

In short: the mechanics reproduce a reading seminar. Each participant prepares their commentary alone (proposal), then everyone listens to the others and formulates disagreements (critique), then everyone adjusts their position (revision). Individual work lays the foundation; subsequent passes refine it collectively. Without the critique phase, it is just a poll with averaging.


Chain-of-Debate vs Graph-of-Debate

CoD is linear and sequential: arguments flow from agent to agent in a fixed order, like links in a chain. This structure is simple to implement but introduces an order bias — arguments formulated last have a greater chance of influencing the consensus.

GoD restructures arguments as a directed graph: each argument is a node, the relationships between arguments (support, refutation, nuance) are edges. The final conclusion corresponds to the best-justified consensus node in the graph, not necessarily the last one produced.

DimensionChain-of-DebateGraph-of-Debate
StructureLinear, sequential roundsNon-linear graph
ArgumentsOrdered flowInterconnected nodes
Order biasPresentReduced
ComplexityLowHigh
ParallelizationPartialStrong

GoD also allows multiple agents to critique the same argument simultaneously, which the chain structure does not handle effectively.


How debate eliminates individual biases

A single LLM is vulnerable to several systematic biases: anchoring on initial hypotheses, overconfidence in domains poorly covered by its training, tendency to produce plausible rather than accurate responses.

The debate mechanism creates external corrective pressure. Work published at ICML 2024 on multi-agent debate (Du et al., 2023) empirically shows that this approach improves factuality and reduces hallucinations on mathematical and strategic reasoning tasks, without modifying the underlying models. Experiments on dialectical knowledge distillation through debate (CoD knowledge distillation) show average improvements of +4.3% on MMMU, +3.8% on MathVista, and +2.6% on CMMMU compared to individual baselines.

Two conditions are necessary for debate to be corrective rather than converging toward the same error:

  • Agent diversity: models must be sufficiently distinct (different prompts, varied temperatures, or different models) so that their errors are not correlated.
  • Substantive critique: agents must produce argued objections, not simply validate the majority.

MASS: automatically optimizing the debate architecture

MASS (Multi-Agent System Search, arXiv 2502.02533) addresses a practical problem: how to optimally configure a multi-agent debate system? The number of agents, their roles, the topology of exchanges, and the prompts form an enormous configuration space.

MASS proposes optimization in three sequential phases:

  1. Block optimization: optimize each agent’s prompts individually before composition.
  2. Topology optimization: search for the best-performing workflow configuration, guided by a measure of each component’s influence.
  3. Global optimization: refine prompts taking into account interdependencies in the complete system.

Published results show an average performance of 78.8% on Gemini 1.5 Pro (8 tasks), surpassing manual architectures by 8+ percentage points, with better token efficiency than simply adding agents.

MASS empirically validates that prompt optimization has more impact than scaling the number of agents — a counter-intuitive conclusion that challenges the “more agents = better results” approach.

Decision matrix: when to use which debate pattern

ContextRecommendationWhy
Simple factual question, latency-criticalDirect answer (no debate)Debate cost = 12×; marginal gain on trivial questions.
Composed reasoning, factuality gain targetedCoD 3 agents × 2 roundsMinimal setup that detects individual biases (+3-5 pts MMMU).
Open problem, multiple competing hypothesesGoD 4-5 agentsArgument graph, reduced order bias, strong parallelization.
New unknown architecture, budget optimizationMASSAutomatic optimization, +8 pts vs manual configurations.
Domain where agents converge (same model, similar prompts)Diversify firstWithout diversity, debate amplifies the common error instead of correcting it.

Strengths and limitations

Strengths:

  • Cross error detection without modifying base models.
  • Documented reduction in hallucinations on factual tasks.
  • Applicable to any existing pair of LLMs.
  • GoD allows formal structuring of arguments, useful for reasoning audits.

Limitations:

  • Computational cost: N agents × R rounds = N×R times the cost of a single inference. For N=4 agents and R=3 rounds, the cost is 12× higher than a direct response.
  • Convergence toward majority consensus: if most agents share the same bias (likely when agents derive from the same base model), the debate amplifies the error instead of correcting it.
  • Order bias (CoD): later arguments influence the conclusion more in a linear structure.
  • Difficult evaluation: measuring the quality of the debate process, not just the final answer, remains an open problem.
  • Superficial critiques: without a mechanism forcing substantive objections, agents tend to validate peers’ responses rather than challenge them (phenomenon documented in “Can LLM Agents Really Debate?”, arXiv 2511.07784).

Key takeaways

  • CoD has multiple independent LLMs debate to mutually correct errors — without modifying the base models.
  • GoD structures this debate as an argument graph, reducing the order bias of sequential chains.
  • The effectiveness of debate depends on the genuine diversity of agents: agents that are too similar converge toward the same errors.
  • MASS automates the optimization of these debate architectures and demonstrates that prompt quality prevails over agent quantity.
  • Computational cost (N×R inferences) is the main practical constraint for production deployment.