In short
A language model generates its answers token by token, with no way to go back. For problems that require several steps — arithmetic, logic, deduction — this architectural constraint is a real issue. In 2022, Google Brain researchers showed that a minimal change to the prompt — including examples with intermediate reasoning steps — was enough to double performance on certain benchmarks. This technique, called Chain-of-Thought (CoT), opened a decade of research on LLM reasoning and remains at the heart of an unresolved scientific debate.
The underlying problem: a model that moves forward without being able to step back
Language models are autoregressive: they predict the next token from all previous tokens. This architecture is powerful for text generation, but it introduces a structural limit for multi-step tasks. An arithmetic or logic problem may require exploring a path, finding it a dead end, and backtracking — something a pure autoregressive model cannot do natively.
Before 2022, LLMs failed in predictable ways on arithmetic benchmarks like GSM8K (8,500 grade-school math word problems in natural language) and logic benchmarks like ARC (Abstraction and Reasoning Corpus). Larger models made progress, but slower than expected on those specific tasks.
Chain-of-Thought: the minimal idea that changed evaluation
In January 2022, Jason Wei and colleagues at Google Brain published “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models” (arXiv:2201.11903). The contribution fits in one sentence: providing a few solution examples that include intermediate steps in the prompt triggers a similar behavior in the model.
Concretely, instead of giving:
Q: Roger has 5 tennis balls. He buys 2 cans of 3. How many does he have? A: 11
You give:
Q: Roger has 5 tennis balls. He buys 2 cans of 3. How many does he have? A: Roger starts with 5 balls. 2 cans of 3 make 6 balls. 5 + 6 = 11.
The model — PaLM 540B in the original experiments — imitates this format and applies the same pattern to new questions. On GSM8K, CoT doubles performance compared to standard prompting and surpasses a GPT-3 specifically trained with an external verifier.
Two empirical observations structure the paper: the effect is absent on models smaller than ~100 billion parameters, and it is robust across three reasoning types — arithmetic, commonsense logic, and symbolic. The authors interpret this as the model exploiting reasoning traces present in large-scale training data.
In short: CoT teaches the model nothing. The mechanism — using generated context as a scratchpad — is already present at training time. The CoT prompt unlocks a latent capability, provided the model is large enough to have internalized it. On small models, the scratchpad does not help — the missing knowledge is not there.
The extensions: self-consistency, least-to-most, Tree of Thoughts
Self-consistency: voting across several paths
Wang et al. (2022, arXiv:2203.11171) push the idea further: instead of generating a single reasoning path, you generate several by sampling, then take a majority vote on the final answer. The intuition is that different paths leading to the same answer constitute a reliability signal. The measured gain is +17.9 points on GSM8K.
Least-to-most: decompose before solving
Zhou et al. (2022, arXiv:2205.10625) propose a prior decomposition: identify sub-problems from easiest to hardest, then solve them sequentially, each solution feeding into the next. On SCAN, a compositional generalization benchmark, GPT-3 reaches 99% accuracy with this method compared to 16% with standard CoT.
Tree of Thoughts: tree-structured reasoning
Yao et al. (2023, arXiv:2305.10601) generalize CoT into a tree-structured framework. The model explores several possible “thoughts” at each step, self-evaluates each branch, and can backtrack. This is an explicit implementation of a tree search (BFS or DFS) over the reasoning space.
Results on the “Game of 24” problem — find a combination of four digits that equals 24 — illustrate the gap: GPT-4 with standard CoT solves 4% of instances; GPT-4 with Tree of Thoughts reaches 74%.
ReAct: reasoning coupled with action
In parallel, the same group published ReAct (arXiv:2210.03629): the model interleaves internal reasoning with calls to external tools — knowledge bases, search engines, APIs. This work lays the foundation for so-called “agentic” architectures, where a model no longer just generates text but interacts with its environment.
In short: self-consistency seeks robustness by voting on multiple attempts. Least-to-most decomposes before solving. Tree of Thoughts allows backtracking. ReAct connects reasoning to the real world. These extensions are not in competition — they cover different problem profiles and often combine.
Which CoT technique for which need
| Context | Recommendation | Why |
|---|---|---|
| Simple arithmetic, short deductive reasoning | Classical CoT | Low cost, typical 2× gain on GSM8K. Sufficient for the majority of cases. |
| Hard math problem, critical reliability | Self-consistency (3-10 votes) | Reduces sampling randomness, ~+18 points on GSM8K. Cost × N attempts. |
| Multi-step planning task (game, combinatorial logic) | Tree of Thoughts | Explores multiple branches, prunes dead ends. Game of 24: 4% CoT → 74% ToT. |
| Question requiring external information | ReAct | The model queries tools mid-reasoning. Foundation of agents. |
| Massive production, tight budget | Zero-shot CoT (“think step by step”) | No few-shot examples to provide, modest but real gain. |
The 2024–2025 shift: training reasoning instead of eliciting it
CoT is a prompting technique — it applies at inference without modifying the model. Starting in 2024, a paradigm shift occurs: train models directly to reason via reinforcement learning, and dynamically allocate more compute to harder problems.
OpenAI o1 (2024) is the first public model to demonstrate this principle: the model “thinks” before answering, and answer quality increases with the thinking time allocated. The training recipe remains confidential.
DeepSeek-R1 (arXiv:2501.12948, January 2025) makes the weights and the method public. Trained via GRPO (Group Relative Policy Optimization) without initial CoT supervision, the model spontaneously develops verification and reconsideration behaviors. On AIME 2024, its score jumps from 15.6% to 71%, reaching 86.7% with majority voting — comparable to o1.
A 2026 paper (arXiv:2602.13517) introduces the notion of “deep-thinking tokens”: tokens where the model’s internal representations shift significantly in the deep layers. The correlation between the proportion of these tokens and final accuracy is r = 0.828, whereas the raw length of the reasoning trace correlates negatively (r = −0.544). In other words, thinking long is not enough — you have to think with intensity.
In short: between 2022 and 2025, reasoning shifts from a prompt technique (free, anyone can apply it) to a property of the trained model (built into weights via RL). “Reasoning” models like o1 and DeepSeek-R1 automatically produce their chain — there is no longer any need to write “think step by step” in the prompt. The cost shifts from the user to inference (more tokens generated).
What the critics show
Fragility under perturbations
Mirzadeh et al. (Apple Research, arXiv:2410.05229, ICLR 2025) publish GSM-Symbolic: the same GSM8K problems with modified names and numerical values. The performance of all tested models degrades. Adding semantically irrelevant information to the question causes drops of up to 65%. The authors conclude that models reproduce reasoning steps present in their training data rather than reasoning formally.
Traces sometimes disconnected from actual decisions
Turpin et al. (arXiv:2305.04388, NeurIPS 2023) show that reasoning chains can be post-hoc rationalizations. The model determines its answer through implicit biases in the prompt, then generates a coherent explanation — one that does not reflect the actual process. Across 13 BIG-Bench Hard tasks, a bias introduced into the prompt degrades accuracy by up to 36% without the model mentioning it in its trace.
Are “emergent abilities” a measurement artifact?
Schaeffer, Miranda, and Koyejo (arXiv:2304.15004, NeurIPS 2023) show that the apparent discontinuity of emergent abilities — including CoT — disappears with continuous metrics. The improvements are linear and predictable; it is the choice of a binary metric that creates the impression of a qualitative jump.
What to remember
- Chain-of-Thought is a prompting technique: including examples with intermediate steps induces similar behavior in large models, without modifying the model.
- The effect is a scale-dependent capability: it is absent or weak on models smaller than ~100 billion parameters.
- The extensions (self-consistency, least-to-most, Tree of Thoughts) each improved results on specific task types, with measured gains on reference benchmarks.
- The shift toward “reasoning models” (o1, DeepSeek-R1) embeds this mechanism into training itself via reinforcement learning, and introduces the concept of test-time compute: allocating more inference compute improves answer quality.
- The question “do LLMs really reason?” remains empirically open: rigorous experiments (GSM-Symbolic, Turpin et al.) reveal structural fragilities incompatible with robust formal reasoning.