In short

An LLM can generate the examples on which it will subsequently be trained. This principle, called synthetic data, made it possible in 2023 to produce models with 7 billion parameters at performance levels comparable to models ten times larger trained on raw data. But the technique has a structural limit: if a model learns only from its own outputs, generation after generation, it eventually irreversibly impoverishes what it can produce. Understanding this mechanism helps read AI laboratory announcements with greater perspective.

In short: imagine a photocopier that copies its own copies. The first reproduction stays faithful. By the tenth generation, the contours fade and rare details disappear. Synthetic data poses the same problem at scale: used as a complement to original data, they enrich; used as a substitute, they silently erode the model’s diversity.


The basic idea: bootstrapping oneself

Training a large language model requires billions of high-quality examples. For a long time, this need was sufficient to justify the massive use of human data — web texts, books, code. But from 2022-2023 onward, an alternative approach took hold: having an existing LLM generate these examples.

The Self-Instruct method (Wang et al., 2022) is the most frequently cited starting point. The principle: provide 175 human-written instructions, let an LLM generate thousands more through imitation and variation, then filter out duplicates and inconsistencies. The result of this loop is used to fine-tune the original model.

Stanford pushed the experiment further with Alpaca (2023): 52,000 examples generated via GPT-3.5, at a cost under 500 dollars. The LLaMA-7B model fine-tuned on this data achieved performance comparable to GPT-3.5 across several evaluations. The idea that human data was irreplaceable took a hit.

In short: with 175 human instructions as a seed and a 500-dollar budget, a team can produce 52,000 training examples of reasonable quality. That is three orders of magnitude below manual annotation costs (typically €5-50 per example). The entry bar for fine-tuning a competent model collapses.


Quality before quantity: the phi hypothesis

Microsoft Research pushed this logic in a surprising direction with the phi series (Gunasekar et al., 2023). Phi-1 is a 1.3-billion-parameter model trained on only 7 billion tokens — including some synthetic code designed to resemble textbooks. Result: 50.6% success on HumanEval (a code generation benchmark), ahead of models ten times larger trained on hundreds of billions of raw tokens.

The hypothesis formulated by the researchers — nicknamed the “textbook hypothesis” — states that the distribution of data matters more than its volume. Text dense in explicit reasoning, well-structured examples, and step-by-step problem solving would teach a model more than an equivalent quantity of ordinary web content.

This thesis remains partially contested. A systematic study published in 2025 (arXiv 2510.01631) shows that synthetic data alone, however carefully curated, does not significantly outperform CommonCrawl — the large corpus of raw web pages — in pre-training. Synthetic data becomes advantageous when mixed with real data, not when it replaces it. The naive promise of “replacing the internet with textbooks” does not hold at scale.


Evol-Instruct: making data harder

Another variant seeks to increase the complexity of synthetic examples rather than their volume. Evol-Instruct (Xu et al., 2023), developed for WizardLM, consists of asking an LLM to rewrite existing instructions to make them harder: adding constraints, introducing multi-step reasoning, shifting the problem to a different domain. This iterative process produces datasets of 250,000 progressively more demanding instruction-response pairs. WizardCoder adapts this strategy to code generation with notable results.


The filtering trap: what we think we are improving

Before even generating synthetic data, teams must clean existing data. Lee et al. (ACL 2022) showed that a 61-word sentence appears more than 60,000 times in C4, a standard training corpus. This excessive repetition leads to memorization rather than generalization.

Deduplication techniques — exact deduplication via suffix arrays, approximate deduplication via MinHash — reduce this problem. Perplexity filtering is another approach: using a small reference model to score the “difficulty” of each document, then discarding texts that are too easy (repetitive boilerplate) or too difficult (incoherent content). It is the best single signal identified to date for data quality (arXiv 2405.20541).

But a 2025 study (arXiv 2510.00866) adds a caveat: common quality classifiers, notably those based on fastText, can create a “quality illusion” — filtered data appears better on standard benchmarks without genuinely improving generalization. The filtering problem is not yet solved.


The curse of recursion

The most structural risk of synthetic data is formulated by Shumailov et al. in a paper published in Nature in 2024 (arXiv 2305.17493): training a model on the outputs of a previous model, repeated multiple times, causes an irreversible collapse of the distribution.

The mechanism is counter-intuitive. It is not the common cases that disappear first, but the rare examples — the distribution tails. After several generations of recursive training, the model loses the ability to produce unusual, nuanced, minority content. It becomes fluent but impoverishing.

The response comes from Gerstgrasser et al. (arXiv 2404.01413, 2024): collapse occurs only if synthetic data replaces real data. If it accumulates alongside real data — without ever supplanting it — collapse is avoided for all tested model types. This accumulation/replacement distinction is the practical rule guiding industrial labs today.

It rests, however, on a fragile assumption: that human data remains available. Work published in 2024 (Villalobos et al., arXiv 2211.04325) estimates that laboratories could reach exhaustion of public human data between 2026 and 2032. If this forecast proves correct, the accumulation/replacement boundary will become difficult to maintain.

In short: the practical rule fits in one sentence — add synthetic data to real data, never substitute. It is the equivalent of a metallic alloy: addition enriches the material, total replacement weakens it. Labs that industrialize LLMs (Meta, Mistral, OpenAI) apply this rule, sometimes implicitly.


Scaling laws and their limits

In 2022, DeepMind published the Chinchilla scaling laws (Hoffmann et al., arXiv 2203.15556): for a fixed compute budget, model size and the number of training tokens must grow proportionally. The empirical rule — 20 tokens per parameter — guided most laboratories’ decisions for two years.

But a replication attempt published in 2024 (Besiroglu & Erdil, arXiv 2404.10102) shows that the coefficients depend heavily on the fitting method: the three approaches in the original paper produce divergent results. The 20-token rule is not a physical constant.

In practice, laboratories such as Meta (LLaMA) and Mistral have deliberately chosen to “over-train” smaller models, well beyond the Chinchilla optimal point. The reason is economic: optimizing long-term inference costs outweighs optimizing training costs. Scaling laws accounting for inference (arXiv 2401.00448) confirm this choice is rational.


Silent benchmark contamination

A transversal bias affects all these evaluations: test data contamination. Several studies (arXiv 2406.04244, arXiv 2406.18824) show that standard benchmarks such as MMLU or HumanEval are partially present in the training corpora of some models. A model can achieve a success rate 4.9 times higher on examples “filtered” from its own training (LessLeak-Bench, arXiv 2502.06215).

The boundary between “the model has learned to reason” and “the model has memorized the answers” becomes indistinguishable without access to training corpora — information that laboratories do not publish. Announced performance on classic benchmarks should be read with caution.


Decision matrix — when to use synthetic data

ContextRecommendationWhy
Domain-specific instruct fine-tuning, limited budgetSelf-Instruct or Evol-Instruct on 10K-50K examplesCost ~€500 vs €5-50 per human example. Comparable performance for fine-tuning on targeted tasks.
Pre-training a foundation modelMix real + synthetic (≤ 30% synthetic)Synthetic alone does not outperform CommonCrawl (arXiv 2510.01631). Mix improves.
Recursive fine-tuning chain (gen N+1 on gen N)Always preserve a fraction of real dataWithout real data: irreversible collapse (Shumailov, Nature 2024).
Complex reasoning tasks to augmentEvol-Instruct (iteratively harder instructions)Improves edge-case coverage without inflating volume.
Final model evaluationPrivate benchmarks, controlled contaminationPublic benchmarks (MMLU, HumanEval) partially contaminated. Score ×4.9 if seen during training.

What to remember

  • Synthetic data makes it possible to massively reduce dependence on human data for fine-tuning: 7B-parameter models fine-tuned on 52,000 synthetic examples have reached the performance of much larger models.
  • Data quality prevails over volume in fine-tuning, but in pre-training, synthetic data alone does not outperform real data — it complements it.
  • Training a model solely on its own outputs, repeated over multiple generations, causes an irreversible collapse of diversity (model collapse). The solution: always retain a base of real data.
  • Chinchilla scaling laws are useful approximations, not constants — over-training small models is often more economically rational.
  • Benchmark contamination is an unresolved structural bias: scores declared by laboratories are difficult to verify without access to training corpora.