In short
Behind every large language model lies a three-act pipeline: first a phase of massive text ingestion across thousands of graphics processors, then a few thousand human examples to learn how to be useful, finally a phase of response comparison to learn how to be preferable. Understanding these three stages means understanding why training a frontier model costs tens to hundreds of millions of dollars — and why a model 100 times smaller can sometimes outperform a giant model after alignment.
Act 1: pretraining — ingesting the internet for months
Pretraining is the foundational stage. The model learns a single task, repeated billions of times: predict the next word in a text. No human correction, no labels — just raw text from the internet, books, scientific articles, and code, in astronomical quantities.
This apparently simple task forces the model to develop an internal representation of language, facts, reasoning, and structures. This is where the model’s deep capabilities reside.
How much data, how much compute?
The question is not trivial. In 2020, OpenAI researchers established scaling laws (Kaplan et al., arXiv:2001.08361): model quality follows a predictable mathematical law as a function of three variables — the number of parameters, the volume of training data, and total compute. The more these three quantities increase, the better the performance, in a regular and quantifiable way.
This framework was corrected in 2022 by DeepMind researchers (Hoffmann et al., arXiv:2203.15556), with what became the Chinchilla law. Their finding: models of the era such as GPT-3 (175 billion parameters) were massively under-fed with data. For a given compute budget, model size and data volume must be scaled equally — approximately 20 training tokens per parameter. Their Chinchilla model (70 billion parameters, 1.4 trillion tokens) outperformed Gopher (280 billion parameters) on the same budget. A result that reconfigured industry strategies.
The Chinchilla law is not, however, the final word. Sardana and Frankle (2024, arXiv:2401.00448) showed that it optimizes training cost but ignores the cost of actual usage: if a model will handle billions of requests, it is better to train it longer to obtain a more compact model, and therefore less costly to run. This is the logic behind Meta’s Llama 3 (arXiv:2407.21783), trained on 15 trillion tokens — well beyond the Chinchilla ratio.
Data curation: the invisible dimension
Ingesting the internet does not mean ingesting everything. The raw web contains a majority of low-quality content: spam, duplicated content, meaningless pages. Filtering pipelines are therefore just as important as the architectures themselves.
HuggingFace documented this process in full with FineWeb (Penedo et al., 2024, arXiv:2406.17557): 15 trillion tokens extracted from 96 web crawls, after deduplication, heuristic filters, and neural classifiers. The variant filtered for educational content produces noticeably better models on reasoning tasks, at equal size. By 2025, synthetic data — that is, texts rephrased or generated by already-capable models — has become a standard ingredient in frontier training pipelines.
What it actually costs
Cottier et al. (2024, arXiv:2405.21015) documented that the cost of frontier models grows by approximately 2.4× per year since 2016. GPT-4 is estimated to have cost between 78 and 100 million dollars in compute alone. Gemini Ultra is estimated at 191 million. Laboratory executives have publicly stated that billion-dollar training runs are imminent.
In short: pretraining is 99% of the cost of a frontier LLM. A GPT-4-class model consumes the equivalent of a small city’s electricity for several months. That is why a few labs (OpenAI, Anthropic, Google, Meta, xAI) dominate the sector — the entry ticket runs in the hundreds of millions of dollars, not counting team and recurring infrastructure.
Act 2: supervised fine-tuning — learning to be useful
After pretraining, the model is capable of coherently continuing a text. But it does not follow instructions: ask it to summarize a document, and it is equally likely to continue the document.
Supervised fine-tuning (SFT, also called instruction tuning) corrects this. The model is provided with thousands of (instruction, ideal response) pairs and trained in the standard way to produce the expected response. Simple in principle, decisive in practice.
OpenAI popularized this approach with InstructGPT (Ouyang et al., 2022, arXiv:2203.02155). A striking result: a 1.3-billion-parameter InstructGPT model, after SFT and alignment, was preferred by human evaluators over GPT-3 with its 175 billion parameters — 100 times more. Raw size is not enough; how the model is guided matters just as much.
The quality of examples matters more than their quantity. Zhou et al. (2023, LIMA) showed that 1,000 carefully selected examples suffice to obtain a highly competitive model. The interpretation: pretraining has already anchored the capabilities. SFT does not create new aptitudes — it unlocks behaviors already present in the model.
In short: SFT is not about “making the model smarter”. It is about “making it obedient to instructions”. A metaphor: a scholar who used to mumble his thoughts aloud becomes capable of politely answering a question. The scholar has not learned more — he has learned to present what he knows in an expected form.
Act 3: preference alignment — learning to be preferable
A useful model is not necessarily a pleasant model to use. Preference alignment is the phase that pushes the model toward responses that humans find better: more accurate, more honest, less harmful, better presented.
The classical RLHF pipeline
Reinforcement Learning from Human Feedback (RLHF) unfolds in three steps.
First, human annotators compare pairs of responses and indicate their preference. Then, a reward model is trained on these comparisons to learn to predict which response would be preferred. Finally, the main model is optimized to maximize this reward score, with a constraint to prevent it from drifting too far from the behavior learned during SFT.
This last constraint is crucial. Without it, the model learns to “cheat”: it discovers that longer or better-formatted responses receive higher scores, regardless of their actual quality. This is reward hacking, a problem documented empirically and actively worked on by several teams (arXiv:2402.09345, arXiv:2502.18770).
DPO: a simpler alternative
In 2023, Rafailov et al. from Stanford (arXiv:2305.18290) proposed a different approach called Direct Preference Optimization (DPO). Their mathematical demonstration: the language model is itself implicitly a reward model. It is therefore possible to derive the optimal policy directly from preference data, without training a separate reward model or resorting to the reinforcement loop. Training reduces to standard classification.
DPO is less costly and more stable than PPO. Meta chose this approach for Llama 3 in production. An important nuance: subsequent work (arXiv:2406.02900) shows that DPO does not escape reward hacking — the degradations observed with RLHF reappear with DPO variants, in different forms.
Constitutional AI: replacing annotators with the model itself
Anthropic proposed a third path with Constitutional AI (Bai et al., 2022, arXiv:2212.08073). The idea: define a “constitution” of ethical principles, then let the model self-critique and self-correct according to these principles. The reward model is trained on preferences generated by the model itself — not by humans. This is what is called RLAIF (Reinforcement Learning from AI Feedback).
The advantage is scalability: the cost of human annotation does not grow with model size. The limitation is inherent to the method: the model remains within its own distribution of biases, without external oversight.
In short: three families of alignment coexist. RLHF = full loop with reward model, classical method. DPO = direct optimization without separate reward model, simpler and more stable. Constitutional AI = the model self-critiques against a written charter. None is universally better; each has its own degradations against reward hacking.
Which training pipeline to choose
| Context | Recommendation | Why |
|---|---|---|
| Generalist frontier model, budget ≥ $50M | Chinchilla-like pretraining + large SFT + full RLHF | Proven recipe, frontier performance, robust alignment. OpenAI/Anthropic pattern. |
| Production model serving billions of requests | Over-training (>20 tokens/param) + SFT + DPO | Smaller model, cheaper to serve, DPO alignment more stable than RLHF. Llama 3 pattern. |
| Specialized model (medical, legal domain) | Base model + targeted SFT + Constitutional AI | Sensitive domain, explicit written charter, alignment scalability without human annotation. |
| POC or academic research | Open model + 1000-10000 example SFT (LIMA-style) | Minimal cost, demonstrates that quality trumps quantity. Sufficient to verify a hypothesis. |
| Language/domain adaptation on existing model | LoRA + targeted SFT | Parameter-efficient adaptation (~1% modified parameters), preserves base capabilities. |
Open debates
Are scaling laws universal? Besiroglu et al. (2024, arXiv:2404.10102) attempted to replicate the Chinchilla results and report problems in the calculation procedures — the published coefficients may be potentially overestimated. Work from 2025 (arXiv:2510.03313) proposes explicitly integrating data quality into the scaling law, something neither Kaplan nor Chinchilla do. The thesis: improving data quality allows reducing model size and compute without performance loss.
Another debate concerns the very necessity of RLHF. LIMA and other works suggest that behavioral alignment comes primarily from pretraining, and that RLHF adds little on models that are well-trained and well-supervised-fine-tuned. This thesis stands in direct contradiction with the massive investments of frontier laboratories in their alignment pipelines.
Key takeaways
- An LLM training pipeline has three distinct stages: pretraining (large-scale text prediction), supervised fine-tuning (instruction-response examples), and preference alignment (RLHF, DPO, or variants).
- The Chinchilla law (DeepMind, 2022) showed that model size and data volume must be scaled equally — but this law ignores inference cost, which pushes industrial practices to depart from it.
- A model 100 times smaller can outperform a giant model after supervised alignment: raw size is not enough.
- Reward hacking is a structural problem of preference alignment: models learn to optimize the proxy (the reward model) rather than actual quality.
- Frontier cost grows by approximately 2.4× per year: GPT-4 is estimated at 78–100 million dollars in compute alone, and billion-dollar training runs are announced as imminent.