In brief

Self-consistency (SC) is an inference strategy introduced by Wang et al. (2023): instead of producing a single response to a prompt, N independent reasoning paths are generated via temperature sampling, then the most frequent answer is selected (majority vote). The method modifies neither the model nor the prompts — it changes only the decoding process. On arithmetic and logical reasoning tasks, gains are documented and reproducible.


Why a single answer is insufficient

A language model generates tokens by selecting the most probable token at each step (greedy decoding). This strategy always produces the same reasoning path — and amplifies the model’s systematic errors: if an intermediate step diverges, the entire chain goes in the wrong direction.

The intuition of Wang et al. is simple: for a reasoning problem, multiple different paths can lead to the same correct answer. If a set of paths is sampled with a non-zero temperature, and the correct answer appears more often than the wrong ones, majority voting selects it. Consistency across paths becomes a reliability signal.


Operational mechanism

SC is applied to a Chain-of-Thought (CoT) prompt (few-shot or zero-shot):

  1. Same prompt, N executions — the prompt remains identical; temperature is raised (typically 0.5–1.0) to introduce diversity.
  2. Path sampling — the model produces N complete reasoning chains, each culminating in a final answer.
  3. Majority voting — final answers are extracted and the one appearing most frequently is retained.

SC is a test-time technique: it requires no additional training and can be applied to any existing LLM.


Quantified gains (Wang et al., 2023)

Across five benchmarks, with identical model and prompt, SC outperforms greedy CoT:

BenchmarkDomainGain
GSM8KArithmetic+17.9 pts
SVAMPArithmetic reasoning+11.0 pts
AQuAAlgebra+12.2 pts
StrategyQACommonsense+6.4 pts
ARC-challengeScience+3.9 pts

These gains are obtained with N = 40 samples. The foundational paper uses PaLM as the reference model.


Computational cost and optimizations

N = 40 means 40 complete inferences per request — a 40× factor on inference cost. This is the main barrier to production adoption.

In short: if one LLM call costs €0.01, an SC response with N=40 costs €0.40. For a single question, this is still acceptable; for 1,000 questions/day in production, you go from €10 to €400/day. The following optimizations (Blend-ASC, adaptive stopping) aim to reduce this inflation to a factor 5-10 rather than 40.

Three directions reduce this cost:

Dynamic allocation (Blend-ASC) — Feng et al. (2025) show that the optimal number of samples depends on question difficulty. Their Blend-ASC method dynamically allocates samples per question and uses on average 6.8× fewer calls than vanilla SC at equivalent accuracy.

Adaptive stopping rule — Cordero-Encinar & Duncan (2025) formalize SC on statistical foundations. They demonstrate that majority voting constitutes a “statistical certificate” of reliability under certain conditions, and introduce a sequential stopping rule (Martingale Majority Certificate) that determines when enough samples have been collected — without iterating to a fixed N.

Repeated sampling at scale — Brown et al. (2024) show that coverage (probability that at least one sample is correct) follows a log-linear power law. On SWE-bench Lite, the resolution rate goes from 15.9% (1 sample) to 56% (250 samples). The first samples provide the most marginal value.


Parallel vs. sequential: a challenge to the paradigm

Wang et al. assume that the N paths are independent — a necessary hypothesis for majority voting to be statistically valid. Sharma & Chopra (2025) directly test this hypothesis by comparing parallel SC (independent paths) to sequential SC (each attempt sees the previous ones).

Result: sequential scaling outperforms parallel SC in 95.6% of tested configurations, with gains up to 46.7%. They also introduce “inverse-entropy weighted voting” — without additional training.

In short: running 40 identical polls in parallel only works if the respondents don’t talk to each other. For an LLM, it is more subtle: its biases are the same at each sample, so 40 independent responses may all converge toward the same wrong answer. Sequential SC (each attempt sees the previous ones) partially corrects this bias.

Implication: if paths are not independent, correlated model errors can be overcounted by majority voting, canceling the benefit of the method. The independence hypothesis is rarely verified empirically. [UNVERIFIED — Sharma & Chopra mention this without direct measurement of error correlation]


SC as a training signal

Beyond decoding, majority voting is used as a ground truth proxy for reinforcement learning without labels. TTRL (Zuo et al., 2025) exploits this signal to train Qwen-2.5-Math-7B solely on unlabeled data: pass@1 on AIME 2024 increases by ~211%.

The idea: if N paths converge on the same answer, that answer serves as a reward signal for training. SC thus bootstraps a fine-tuning process that makes the model more reliable at inference — ultimately reducing the need for SC itself.


Limitations

Tasks with a single verifiable answer — SC works well on problems where a single correct answer exists (arithmetic, formal logic). For open-ended generation tasks or those with multiple valid answers, majority voting has no clear meaning.

Systematic biases — If the model has a strong bias toward a type of error, the N paths can converge on the same wrong answer. SC then amplifies the error instead of correcting it.

Production cost — The literature evaluates SC primarily on large models (PaLM 540B, GPT-4). Effectiveness on 7B–13B parameter models is less documented. [UNVERIFIED — few SC benchmarks on small models published]

Applicability to long tasks — SC aggregates final answers, not intermediate reasoning steps. For multi-step chains where an early error propagates, SC does not detect the error in the path — it votes on answers that may all be wrong.


When to use SC versus the alternatives

ContextRecommendationWhy
Verifiable arithmetic/logic problem, acceptable costSelf-Consistency, N=20-40+6 to +18 points on GSM8K/SVAMP/AQuA, robust, simple
High-volume production, limited token budgetBlend-ASCDynamic allocation, 6.8× fewer calls at equivalent accuracy
Exploratory task where planning mattersTree of Thoughts (ToT)Explicit backtracking, 74% versus 4% for CoT on Game of 24
Iterative solving with learning across attemptsSequential SC+46.7% over parallel SC on 95.6% of configurations (Sharma 2025)
Open-ended task (writing, brainstorming)No SCThere is no valid “majority answer” — voting without a signal

Relationship to other techniques

SC fits within a hierarchy of reasoning techniques (Gulli, 2025, App. A):

  • Linear CoT — a single exposed reasoning path; the basis for SC.
  • Self-Consistency — multiple parallel CoT paths + majority vote.
  • Tree of Thoughts (ToT) — tree-based exploration with explicit backtracking; preferable when planning and solution space exploration matter. On Game of 24: 74% (ToT) vs 4% (standard CoT) with GPT-4.

SC is a decoding strategy; ToT is a search strategy. Both build on CoT but address different problem profiles.


Key takeaways

  • Self-consistency generates N independent reasoning paths and retains the majority answer — without modifying the model or the prompt.
  • Gains are documented on verifiable reasoning tasks: +6 to +18 points depending on the benchmark (Wang et al., 2023, ICLR).
  • Cost is the main barrier: N = 40 means 40 calls per request. Dynamic allocation methods (Blend-ASC) reduce this factor by 6.8×.
  • The path independence hypothesis is challenged: sequential SC outperforms parallel SC in the majority of tested configurations.
  • SC does not apply to open-ended tasks or long chains where errors propagate between steps.