In Brief

Four techniques structure LLM reasoning today: Chain-of-Thought (linear decomposition), Tree-of-Thought (tree-based exploration), ReAct (reasoning interleaved with actions), and the Scaling Inference Law (allocating more compute to inference rather than increasing model size). These approaches are not mutually exclusive — they cover distinct use cases and combine in advanced systems. A smaller model that “thinks longer” can outperform a larger model that answers directly.


Why These Techniques Exist

By default, an LLM produces a response token by token, without explicating its intermediate steps. For simple questions, this direct generation is sufficient. For problems requiring multiple logical steps — computation, planning, diagnosis — it frequently fails.

The reason is structural: the model has no “scratch pad.” All its reasoning must fit within the probability of the next token. Reasoning techniques work around this limitation by forcing the externalization of intermediate steps into the generated sequence.

The goal is not to make models “smarter” in some vague sense. It is to structure the generation process so that each step benefits from previous steps — exactly like a human who writes out their calculations rather than doing them in their head.

In short: an LLM without a reasoning technique is a human asked to compute 47 × 83 in their head, while speaking. With a reasoning technique, it is the same human taking a scratchpad. The scratchpad is not magic — it simply offers external memory the model can re-read while generating its next token.


Chain-of-Thought (CoT) — Explicit Linear Reasoning

Principle

Chain-of-Thought consists of asking the model to generate its intermediate steps before producing its final answer. The best-known formulation is “think step by step.” In practice, you can also provide few-shot examples showing the expected reasoning structure.

Each step in the chain decomposes the original problem into a simpler sub-problem. The final answer is only produced once the complete chain is generated.

What It Changes

CoT significantly improves performance on tasks requiring multiple chained inferences: arithmetic, logical reasoning, multi-step deduction. Explicit decomposition reduces consistency errors — the model can “verify” its own steps by reading them in context.

Limitations

CoT is linear: it follows a single reasoning path. If it takes a wrong direction in the early steps, there is no backtracking mechanism. For problems where multiple solution approaches exist and some are dead ends, this linearity is a handicap.

CoT also operates entirely on the model’s internal representations — it cannot query external sources or update its knowledge mid-chain.

In short: CoT is the detective who follows their hunch all the way without ever asking if they got the wrong hypothesis. As long as the initial hunch is right, it works very well and is fast. When it is wrong, all the time spent reasoning is lost.


Tree-of-Thought (ToT) — Tree-Based Exploration

Principle

Tree-of-Thought extends CoT by allowing the exploration of multiple reasoning paths in parallel, organized as a tree. At each node, the model generates several candidate branches, evaluates their viability, and prunes non-promising paths before continuing.

The tree can be traversed depth-first (exploring one path to its conclusion before testing another) or breadth-first (evaluating all candidates at one level before moving to the next). Heuristics — often the model itself acting as an evaluator — guide the pruning.

What It Changes

ToT is designed for problems where the solution requires strategic backtracking: planning, strategy games, differential diagnosis, code generation with multiple constraints. It allows dead ends to be detected before investing the entire generation budget in them.

Empirically, ToT outperforms CoT on planning and combinatorial reasoning benchmarks where direct paths frequently fail.

Limitations

ToT is expensive in tokens and model calls. Evaluating and pruning each branch multiplies the number of inferences. For simple or well-structured problems, this cost is not justified — CoT is sufficient and less expensive.

The quality of pruning also depends on the quality of the evaluation heuristic. A poor evaluator prunes good branches and retains bad ones.

In short: ToT multiplies cost by 5 to 20× depending on tree depth and branch width explored. On Game of 24-type tasks, the gain (4% → 74%) justifies the cost. On a simple question, it is using a jackhammer to drive a nail.


ReAct — Reasoning Interleaved with Action

Principle

ReAct (Reasoning + Acting) breaks the boundary between internal reasoning and environment interaction. The model operates in an iterative loop: Thought → Action → Observation → Thought…

  • Thought: the model explicates its reasoning about the current problem state.
  • Action: the model invokes an external tool (web search, database, calculator, API).
  • Observation: the result of the action is injected into the context.
  • The cycle repeats until the task is resolved.

The key is interleaving: thought informs action, and observation modifies the next thought. Reasoning is no longer cut off from the real world.

What It Changes

ReAct is the operational foundation of autonomous agents. Where CoT and ToT remain confined to the model’s context, ReAct enables acting on the environment and receiving feedback from it. A ReAct agent can correct an error in real time: if a search does not return the expected results, the next thought integrates this and adjusts the action.

Evaluations on multi-step question-answering benchmarks (HotPotQA, Fever) show that ReAct outperforms pure action models while remaining competitive with CoT alone — and ReAct+CoT combinations outperform both taken separately.

Limitations

ReAct requires accessible and reliable tools. The loop can derail if returned observations are noisy or contradictory. Without an explicit termination mechanism, the model can enter non-converging loops.

Latency accumulates with each tool call — a task resolved in 8 iterations takes the time of 8 successive calls.

In short: ReAct is the foundation of modern agents. It is not suitable for instant-answer questions — it is a thought/action cycle that takes seconds or even minutes depending on the tools invoked. But it makes possible tasks an LLM alone cannot do: searching in real time, executing code, manipulating files.


Scaling Inference Law — Thinking Longer Beats a Larger Model

Principle

The Scaling Inference Law (also called test-time compute scaling) postulates that model performance improves predictably with computational resources allocated to inference — not training. In other words: the same model, given more “thinking budget,” produces better responses.

OpenAI materialized this principle with o1 (September 2024): the model generates internal reasoning tokens (thinking tokens) before producing its visible response. These tokens constitute a structured draft, invisible to the user but central to output quality.

What It Changes

The implications are concrete. On the AIME 2024 benchmark (competition mathematics), GPT-4 solved approximately 9% of problems, versus 79% for o1 with its extended reasoning budget. Subsequent work (DeepSeek-R1, 2025; s1, 2025) confirms that smaller models, trained with reinforcement on verifiable responses (RLVR — Reinforcement Learning with Verifiable Rewards), catch up with and even surpass larger models on formalized reasoning tasks.

The Scaling Inference Law redefines the cost/performance ratio: rather than investing in a larger model, you can invest in more tokens at inference time, which is often cheaper and more flexible.

Limitations

The gain is most visible on tasks with verifiable answers: mathematics, code, formal logic. On open-ended tasks (creative writing, strategic advice), the effect is less pronounced — the absence of a verification signal limits the reinforcement learning that underlies reasoning models.

Generating more reasoning tokens costs more at inference. The budget must be calibrated to the task type: using o1 to write a simple email is like using a jackhammer to drive a nail.

In short: the Scaling Inference Law shifts the economics. Instead of investing in a 10× larger model (training cost), you pay for 10× more thinking tokens at inference (marginal cost). The benefit concentrates on domains where a “good answer” is measurable — for creativity, reasoning longer no more helps than a human ruminating on a poem.


Comparative Table — When to Use Each

TechniqueStructureStrengthsLimitationsTypical use cases
CoTLinearSimple, efficient, low costNo backtracking, no external accessArithmetic, logical deduction, explanation
ToTTree-basedMulti-path exploration, backtrackingHigh cost, depends on evaluatorPlanning, diagnosis, combinatorial puzzles
ReActIterative loopReal-time environment interactionCumulative latency, loop riskAutonomous agents, information retrieval, multi-tool tasks
Scaling InferenceToken budgetStrong gains on formalized tasks, flexibleInference cost, limited to verifiable tasksMathematics, code, formal reasoning

Combinations Observed in Practice

These techniques are not mutually exclusive. Modern implementations combine them:

  • CoT + ReAct: the model reasons explicitly (CoT) at each turn of the ReAct loop. This is the base pattern of most current agents.
  • ToT + Scaling Inference: allocating more tokens to tree-based exploration improves the quality of retained paths.
  • ReAct + Scaling Inference: o1 and DeepSeek-R1 models use internal reasoning tokens before each action, which amounts to integrating Scaling Inference into each ReAct iteration.

Google DeepResearch (2025) illustrates the composition: initial query generation → searches → gap analysis → refined queries → final synthesis. This orchestration implicitly uses CoT for planning, ReAct for search engine interactions, and allocates a larger reasoning budget to synthesis steps.


Key Takeaways

  • CoT: forcing the model to write its intermediate steps improves performance on any decomposable problem. Low cost, widely adopted.
  • ToT: useful when multiple solution paths exist and some are dead ends — requires a branch evaluation mechanism.
  • ReAct: the fundamental pattern for autonomous agents. Reasoning and action feed each other at each iteration.
  • Scaling Inference Law: thinking longer with the same model can be worth a larger model. Gain is maximal on tasks with verifiable answers (math, code).
  • These techniques combine: the highest-performing systems use CoT or Scaling Inference inside ReAct loops.