In short
Doubling the size of a model improves its performance in a predictable way. This is not an intuition: it is a mathematical law, verified over seven orders of magnitude. Researchers call it a scaling law. It has guided the development of GPT-3, and then of almost every large model that followed.
But this law has surprising corollaries, concrete limits, and since 2024 a serious competitor: inference-time scaling.
Scaling laws answer a simple question: if I invest more — more parameters, more data, more compute — how much performance do I gain?
The answer is regular. Too regular to be a coincidence.
Kaplan’s discovery (2020)
In January 2020, OpenAI’s team publishes a paper that will structure a full decade of research. Jared Kaplan and his colleagues train dozens of models of varying sizes and measure their performance with a simple metric: cross-entropy loss — the error in predicting the next token.
The result is clean. Performance follows a power law with three variables:
- N: the number of model parameters
- D: the number of tokens seen during training
- C: the total compute budget (in FLOPs)
Each doubling of one of these variables produces a regular, predictable performance gain. The relationship holds over seven orders of magnitude — the gap between a model of a few million parameters and one of several hundred billion.
This result has an immediate consequence. For a fixed compute budget, it is better to train a large model on a small amount of data than a small model on a large amount of data. GPT-3 (175 billion parameters) is sized directly from this principle.
In short: Kaplan’s law is the observation that LLM performance does not progress in unpredictable jumps but along a regular curve tied to three levers (parameters, data, compute). Knowing three points on this curve, you can predict the fourth. This regularity justified hundreds of millions of dollars of investment: people roughly knew what they were buying.
The Chinchilla correction (2022)
Two years later, DeepMind publishes a revision that overturns Kaplan’s prescription on a fundamental point.
Jordan Hoffmann’s team trains more than 400 models, this time varying both model size and data volume simultaneously. Their conclusion: Kaplan was wrong on a specific point. Model size and token count must grow proportionally. For each doubling of parameters, you should double the training tokens.
The rule that emerges is called the Chinchilla rule: 20 tokens per parameter. A 70-billion-parameter model should see 1.4 trillion tokens to be trained optimally.
To demonstrate the principle, DeepMind trains Chinchilla — 70 billion parameters on 1.4 trillion tokens. Result: Chinchilla beats Gopher (280B), GPT-3 (175B) and Megatron-Turing NLG (530B) on most benchmarks. With four times fewer parameters than Gopher.
The implication is awkward for the industry: the large models of 2021-2022 were massively under-trained on data. Too much had been invested in parameters, not enough in tokens. The community realigned its practices accordingly.
In short: Kaplan said “increase parameters above all”. Chinchilla corrected this by saying “increase parameters and data together, at 20 tokens per parameter”. Direct consequence: a 70-billion-parameter model well-trained (1,400 billion tokens) beats a 280-billion poorly-fed model. Size alone does not make performance.
Parameters are not everything
Kaplan’s and Chinchilla’s scaling laws measure loss — the error in predicting the next token. That is not the same thing as a model’s real-world capabilities on concrete tasks.
This distinction matters. In 2022, Jason Wei and his Google colleagues describe emergent capabilities: abilities absent below a certain parameter threshold that seem to appear abruptly beyond it. Solving multi-step equations, transliterating, performing modular arithmetic.
The counter-argument comes the following year. Rylan Schaeffer and his co-authors (NeurIPS 2023, best paper) show that these jumps are partly measurement artifacts. When you use a discontinuous metric — like exact-match success rate — continuous progress in loss translates into an apparent jump in the numbers. With continuous metrics, progression is often smooth.
The debate is not closed. On certain tasks, jumps persist even with continuous metrics. But the idea that scaling produces spectacular, unpredictable emergences is now contested with solid arguments.
The data wall
Scaling laws assume an unlimited supply of quality data. But that stock is finite.
Epoch AI estimates that public human text of good quality, once duplicates are stripped, represents about 300 trillion tokens. At accelerating training rates, this capital could be exhausted between 2026 and 2032.
Facing this constraint, two strategies emerge. First strategy: repeat the data. A study by Muennighoff et al. (NeurIPS 2023) shows that repeating up to four training epochs does not significantly degrade performance. Beyond that, returns collapse.
Second strategy: deliberately over-parameterize. Recent models like Llama 3 are trained well beyond the Chinchilla prescription — more tokens than compute-optimality requires. The idea is to maximize inference performance, at the cost of compute efficiency during training. A smaller but more-trained model is cheaper to deploy.
A third dimension: inference-time compute
September 2024. OpenAI releases o1. The model is distinguished not only by its size or training data — it thinks longer before answering, generating internal chains of thought via reinforcement learning.
Performance on math problems grows log-linearly with the compute time allocated at inference. The more the model is allowed to “think”, the better it performs. o3, released in December 2024, pushes the logic further: 88% on GPQA, 88% on ARC-AGI, 25% on FrontierMath.
A paper by Snell et al. (2024) formalizes this principle: for reasoning tasks, inference-time scaling can be more compute-efficient than training a bigger model. In other words, spending compute at answer time can be worth more than spending it during training.
This finding opens a new scaling dimension. But it has a cost: o3 in high-performance mode can cost over 1,000 dollars per task. The gain is real, but its accessibility is bounded by economics as much as by technique.
In short: since 2024, the question is no longer just “how much to spend at training”, but also “how much to spend per query”. For tasks like solving math problems, paying 100× more tokens at inference with a mid-size model can outperform a 100× larger classically-trained model. It is a major economic shift.
Which dimension to scale by objective
| Context | Recommendation | Why |
|---|---|---|
| You are training a base model at fixed compute cost | Follow Chinchilla (20 tokens/param) | Maximizes performance per dollar at training, demonstrated on 400 models. |
| You want to minimize inference cost (millions of queries) | Over-train (beyond Chinchilla, 75-200 tokens/param) | Smaller model at equivalent performance — cheaper to serve. Llama 3 made this choice. |
| You are solving hard one-off problems (math, code) | Reasoning model (o1, o3, R1) with high token budget | Test-time compute scaling: 100× more tokens at inference on verifiable tasks = big gains. |
| You are constrained by quality data availability | Reduce model size, increase data quality + epochs | Densing law: at equivalent performance, 2025 models are 30-100× smaller than 2022. |
Capability density
Alongside these debates, a deeper trend is emerging. Xiao et al. (2024, Nature Machine Intelligence) measure not raw model performance, but capability density: performance divided by parameter count.
Since 2023, this density has doubled roughly every 3.5 months. At equivalent performance, recent models need exponentially fewer parameters than their predecessors. Sparse architectures (MoE — Mixture of Experts) and data quality are the main drivers.
This result suggests that the “scaling wall” may not be measuring the right thing. If you are looking for bigger models, the wall is real. If you are looking for more capable models, progression continues — in a different form.
What to remember
- Scaling laws establish that LLM performance scales predictably with model size, data volume and compute. This regularity, verified over seven orders of magnitude, has guided the development of large models since 2020.
- Kaplan (2020) set the foundation: more parameters, at fixed compute budget, beats more tokens. Chinchilla (2022) corrected this bias: size and data must grow together, at a rate of 20 tokens per parameter. The community massively readjusted its practices.
- Since 2024, a new dimension has been added: inference-time compute. o1 and o3 show that letting a model “reason longer” improves performance predictably, independently of parameter scaling.
- The limits are clear: the stock of quality human data is finite, classic benchmarks are saturating, and test-time compute costs remain prohibitive for many use cases. The debate on “emergent capabilities” — spectacular jumps tied to size — is active and unresolved.
- The question is no longer “does scaling work?” but “which dimension to scale, with what data, for what goal?”