In Brief
Models like DeepSeek-R1 and QwQ-32B, trained by reinforcement on reasoning tasks, spontaneously develop conversational structures in their internal traces: questioning, perspective shifts, conflict between positions, then reconciliation. These behaviors were never explicitly requested. A study published in January 2026 establishes a causal link between these patterns and accuracy gains on the most demanding benchmarks.
A Distinction to Establish from the Outset
Unlike chain-of-debate, where debate is orchestrated by the prompt — multiple model instances explicitly respond to each other over several turns — reasoning models develop these conversational structures spontaneously under reinforcement learning training, inside a single reasoning trace, without any instruction requesting it.
This is not a prompt engineering technique. It is an emergent phenomenon observed after RL training, whose precise mechanisms remain to be characterized.
What Kim et al. (2026) Observe
The study “Reasoning Models Generate Societies of Thought” (Kim, Lai, Scherrer, Agüera y Arcas and Evans, arXiv:2601.10825) analyzes 8,262 problems drawn from six high-level benchmarks — BigBench Hard, GPQA, MATH Hard, MMLU-Pro, MUSR, IFEval — using an LLM-as-judge to detect conversational patterns in the reasoning traces of DeepSeek-R1 and QwQ-32B.
Four structures emerge in a statistically distinct manner compared to instruction-tuned models of comparable size:
- Internal Question-Answering: the model poses questions to itself and answers them (β=0.345, p<10⁻³²³ for DeepSeek-R1).
- Perspective shifts: explicit viewpoint changes appear in the trace (β=0.213).
- Conflict of perspectives: disagreements are formulated between distinct internal “voices.”
- Reconciliation: these conflicting positions converge toward a unified answer.
What is striking in these figures: all instruction-tuned models tested, from 8 to 671 billion parameters, show a low prevalence of these behaviors. Model size is not sufficient to produce them — it is the type of training that matters.
In short: a classical large model (instruction-tuned) responds directly, like a lone speaker on the rostrum. A model trained by reinforcement on hard problems becomes an assembly — it questions, contradicts, adjusts itself. This difference is not in size but in training method: it is RL on a binary reward signal that produces the “internal society”, not adding parameters.
The measurement relies on an LLM-as-judge annotated to detect these structures in the free text of reasoning traces. This method introduces a potential source of bias: the judge can itself produce classification artifacts. Kim et al. control for this effect by comparing distinct model families, which makes systematic bias unlikely — but the method is not without limitation. Manual annotation of a subsample would have constituted stronger validation; the authors do not explicitly mention it as having been performed.
The Causal Link with Performance
Showing that patterns exist in traces does not prove they cause performance — they might simply be correlated with other model properties. Kim et al. address this question through structural equation modeling.
The central result: conversational steering predicts accuracy with β=0.228 (p<10⁻²²), mediated by four measurable cognitive strategies — step verification (Δ=5.815), backtracking (Δ=0.881), sub-goal decomposition (Δ=0.621), backward reasoning (Δ=0.809).
The most direct demonstration uses a Qwen-2.5-3B model trained by PPO on the Countdown task. By applying positive conversational steering (+10), accuracy goes from 27.1% to 54.8%. Negative steering (−10) brings it back to 23.8%. This is not a correlation observed after the fact: it is a causal manipulation.
In short: steering is a lever you can turn one way or the other. If you amplify the model’s internal conversational patterns, its accuracy nearly doubles (27% → 55%). If you suppress them, it drops. Patterns make performance — not the other way around.
The authors also note a diversity of “personalities” in the traces — agreeableness diversity (β=0.297) and neuroticism (β=0.567) for DeepSeek-R1 — suggesting that internal voices are not homogeneous. They draw, with caution, on the work of Minsky (The Society of Mind, 1988) and theories of the “dialogic self” (Bakhtin, Mead), without claiming a direct neural correspondence.
Why RL Without Supervision Produces These Structures
The underlying question is: why does RL training on a binary reward signal (correct answer or not) make social structures emerge in reasoning?
DeepSeek-AI documents the phenomenon as early as the original DeepSeek-R1 paper (arXiv:2501.12948, published in Nature in 2025). Training uses GRPO — Group Relative Policy Optimization — an algorithm that removes PPO’s value model and estimates advantages directly from the intra-group distribution. The reward signal is entirely rule-based: correct format + correct answer. No supervision on the quality of intermediate reasoning.
Yet non-rewarded behaviors emerge: self-reflection, step verification, dynamic strategy adaptation, and — documented by the authors — an “aha moment” where an intermediate model spontaneously develops the ability to reconsider its own reasoning in an anthropomorphic tone. The performance speaks for itself: on AIME 2024, DeepSeek-R1 goes from 15.6% (zero-shot) to 71.0% (pass@1) and 86.7% with majority voting, at the level of OpenAI-o1 — without these emergent behaviors being targeted by the reward signal.
In short: nobody told the model “debate with yourself”. They just told it “answer correctly”. But the most efficient path to that reward goes through doubt, verification, backtracking. The model discovered these adaptive strategies because they work, not because they were shown.
The intuition behind this result: a model optimized to find correct answers to difficult problems must, at some point, learn to doubt its own intermediate solutions. Verification, backtracking, reformulation from a different angle — these are rational adaptive behaviors in the face of problems that resist a direct approach. The model discovers them not because they were shown to it, but because they constitute the path to the reward.
This phenomenon fits within a broader framework, the RLVR paradigm (Reinforcement Learning with Verifiable Rewards): objective and verifiable reward signals — correct answers, unit tests, formal proofs — trigger the emergence of new solution patterns that no one explicitly programmed. Code-based reasoning on mathematical tasks goes from 65% to over 90% after RLVR without specific instruction. Med-RLVR produces step-by-step clinical reasoning in medical domains without a single reasoning annotation.
The internal conversational structure appears to be a solution the model discovers because it works — not because it was designed.
A Theoretical Root: AI Safety via Debate (2018)
The idea that adversarial structure improves the epistemic quality of a system has an older formal root. Irving, Christiano, and Amodei (arXiv:1805.00899, 2018) propose training agents through self-play on a zero-sum debate game: two agents debate in turns before a human judge. The theoretical result is solid — with optimal play, debate can answer all questions in PSPACE with polynomial-time judges, versus NP only for direct judgment.
This framework operates with explicitly adversarial external agents, trained for that purpose. What Kim et al. show is of a different nature: debate structures instantiate inside a single model, without a game, without an adversary, without instruction. The form is comparable; the mechanism is distinct.
What This Changes Relative to External Multi-Agent Debate
The research line on Multi-Agent Debate (Du, Li, Torralba, Tenenbaum, and Mordatch, arXiv:2305.14325, ICML 2024) had shown as early as 2023 that multiple LLM instances debating over several turns improve factuality and reasoning. This demonstrates that collective structure helps — but it operates on models not trained to debate, and organizes debate at the prompt level.
The question posed by Kim et al. is different: do these structures emerge spontaneously inside a single model, without external orchestration? The answer appears to be yes, but with an important caveat: in 2026, no paper directly compares, on the same tasks and with the same computational budget, internal emergence and external MAD frameworks. The two lines progress in parallel without an experimental bridge.
There are also known limitations on the external MAD side: Wang et al. (ACL 2024, arXiv:2402.18272) show that a single-agent with a strong prompt achieves the same performance as the best MAD frameworks on a wide range of tasks. The ICLR 2025 analysis confirms that homogeneous MAD frameworks frequently convert correct answers to incorrect ones, and that Self-Consistency outperforms most MAD approaches on nine benchmarks. Only heterogeneous cooperation — mixing different models — shows consistent gains.
The internal “Society of Thought” escapes these specific criticisms, because it is not a prompting technique. But it raises its own open questions.
A complementary perspective comes from Riedl (arXiv:2510.05174, 2025), who measures whether a multi-agent LLM system forms an integrated collective or a simple aggregation of independent instances. His method — partial information decomposition of time-delayed mutual information — reveals that differentiated personas combined with instructions for modeling other agents produce goal-oriented complementarity. This result comes from the external; the question of whether the “internal society” of RL models exhibits the same informational integration properties remains unanswered.
Limitations and Open Questions
Causality or epiphenomenon? Kim et al. establish a causal link through structural equation modeling, but causal identifiability is not complete. Conversational patterns might be a consequence of the same parameters that improve accuracy, rather than a distinct cause. The paper does not directly refute this objection.
Unknown neural mechanism. The correspondence between structures observed in reasoning traces and the internal circuits of the transformer — which attention heads, which circuits activate these distinct “voices” — is not established. This is a gap acknowledged by the authors as a future direction.
Generalization outside formal domains. The results concern mathematical and formal reasoning tasks, where a binary reward is possible. Open-ended domains — creative writing, nuanced argumentation — remain outside the RLVR paradigm’s reach. It is unknown whether conversational emergence reproduces in these contexts.
Replicability. Kim et al.’s results rest on DeepSeek-R1 and QwQ-32B. The presence of “Society of Thought” structures in other model families (Llama, Mistral, GPT-4o under RL) is not documented to date. DeepSeek-R1 and QwQ-32B are both from Asian labs with similar training processes — it remains plausible that model families trained differently do or do not produce these structures.
What this phenomenon does not say. The emergence of internal conversational structures must not be confused with the presence of “sub-minds” or distinct cognitive entities inside the transformer. Kim et al. observe patterns in text traces, measured by a classifier. These patterns are causally linked to performance — that is the central and solid result. But the interpretation in terms of “society” or “personalities” is a useful metaphor for analysis, not a description of the model’s internal architecture.
Key Takeaways
- DeepSeek-R1 and QwQ-32B spontaneously develop, under RL training without reasoning supervision, internal conversational structures: questioning, disagreement, reconciliation.
- These patterns are not prompted — they emerge because they improve performance on binary reward signals.
- A causal link has been established: steering these structures modifies accuracy in a direct and measurable way.
- Model size is not sufficient to produce these behaviors; it is the type of training (RL on verifiable rewards) that triggers them.
- The underlying neural mechanism remains uncharacterized, and generalization outside formal domains is an open question.