In short

Model distillation is a training technique: a large model, called the teacher, guides the learning of a small model, the student. The goal is not to copy the teacher’s parameters — it is to transfer what it knows. The result is a compact, fast model, deployable where the teacher is not.

In short: a 70-billion-parameter model weighs 140 GB at full precision — unthinkable on a phone. Distillation produces a condensed 7-billion-parameter version that often retains 85-95% of capabilities. It is the equivalent of a great chef training a sous-chef: you do not give them the restaurant, you transmit the techniques. The sous-chef will never be as good, but they will do 80% of the work for 10% of the cost.


The problem distillation solves

Large language models (GPT-4, Llama 70B, DeepSeek R1) are powerful but costly: in energy, memory, and latency. Deploying them on a phone or in an embedded system is impractical. The obvious solution — training a smaller model directly — works, but often produces a mediocre model. The small model lacks enough signal to learn what the large one extracted from billions of data points.

Distillation changes the problem. Rather than learning from the raw world, the student learns from the teacher. It benefits from a signal that is already distilled, organized, structured.

The master and the apprentice

The analogy holds. A master craftsman does not teach by showing thousands of failed pieces — he shows how he reasons when faced with a problem. He says: “this material is like that one, it is worked in such a way.” The apprentice learns faster because the teaching is concentrated.

In machine learning, this “how he reasons” translates into soft targets: the full probability distribution produced by the teacher on each example. When the teacher predicts a token, it does not assign 100% to a single candidate. It distributes its confidence — 60% for “cat”, 30% for “feline”, 5% for “tiger”, etc. This distribution carries information about relationships between concepts. It says that “cat” and “tiger” are closer than “cat” and “table”.

A classic binary label (hard target) erases this information. Distillation preserves it.

The concept was formalized by Hinton, Vinyals and Dean in 2015. The distillation loss function combines two signals: cross-entropy on true labels, and divergence between teacher and student distributions. A temperature parameter (T > 1) softens distributions to amplify weak signals.

In short: the magic lies in the “soft targets”. When the teacher hesitates at 60% “cat” / 30% “feline” / 5% “tiger”, it delivers a map of conceptual proximity that no binary label contains. The student learns the geography of the domain in addition to the right answer — which explains its ability to generalize despite its reduced size.

Three major families of methods

Logit distillation

This is the original form, described by Hinton in 2015. The student imitates the teacher’s outputs — its logits or probabilities — without access to internal layers. Simple to implement, applicable in a black-box setting.

DistilBERT (Hugging Face, 2019) is the canonical example. It is 40% smaller than BERT, 60% faster, and retains 97% of GLUE performance. It combines logit distillation, alignment of hidden representations, and standard training on data.

Intermediate representation distillation

Here, the student learns to reproduce not only the teacher’s outputs, but also its internal states: hidden layer activations, attention matrices. This additional signal guides compression more precisely. The trade-off is that it requires “white-box” access to the teacher — its weights must be accessible, not just its outputs.

Data distillation

A lesser-known but effective variant. The teacher is not a direct supervision signal — it is a data generator. It is asked to produce high-quality, carefully filtered examples that will be used to train the student. LIMA (2023) illustrates this principle: 1,000 well-crafted examples are enough to approach the teacher’s performance on many tasks.

Microsoft Phi-4 (2025, 14B parameters) pushes this principle far. It combines synthetic data distillation and intensive quality filtering. The result is a model that rivals much larger models on mathematics, code, and reasoning benchmarks — and that runs locally on CPU.

Reasoning distillation: the qualitative leap

Classic distillation transfers competence: the student learns to produce outputs close to the teacher. Reasoning distillation goes further — it attempts to transfer the thinking process itself.

Microsoft Orca (2023) is the first significant example. The teacher’s prompting (GPT-4) includes an explicit instruction: “reason step by step and justify your answer.” The reasoning traces produced — not just the answers — serve as supervision for the student (13B parameters). Orca outperforms models that are nevertheless larger on complex reasoning benchmarks.

DeepSeek-R1 (2025) pushes this principle to industrial scale. DeepSeek distills its reasoning model (671B parameters) into dense models ranging from 1.5B to 70B parameters. The 32B distilled model outperforms OpenAI o1-mini on several benchmarks. What is being transferred here are chains of thought — qualitative capabilities, not simply compression.

Distillation, quantization, pruning: three complementary techniques

These three approaches complement each other. Confusing them leads to poor architecture decisions.

Pruning: redundant parameters are removed from the model — weakly active connections, unnecessary attention heads. The architecture becomes lighter. Accuracy drops moderately (approximately 89% retained in recent benchmarks).

Distillation: the lightened architecture is retrained under the teacher’s supervision. It recovers some of the capabilities lost during pruning. It is the only technique capable of recovering performance after compression. It is also the only one that enables the transfer of qualitative capabilities.

Quantization: the numerical precision of weights is reduced (float32 → int8 or int4). The model takes less memory, computes faster. No architectural modification.

The empirically identified optimal sequence: pruning → distillation → quantization. Pruning lightens the structure. Distillation recovers quality. Quantization finalizes compression without disturbing the structural changes already made. Order matters: applying quantization before distillation would degrade the supervision signal.

A full pipeline can take a model of 2.7B parameters in float32 to an equivalent model of 700M in int8 — a reduction of approximately 14× — while retaining 85% of the teacher’s performance.

In short: confusing these three techniques is the most frequent mistake. Quantization compresses the numbers (16 bits → 4 bits, ×4 memory gain). Pruning removes connections (compute gain but ~10% precision loss). Distillation retrains a new model from the teacher’s “soft labels”. Only distillation can recover lost performance. The three stack: pruning first (structure), distillation next (quality), quantization as finishing (memory). Reversing the order degrades the result.

Decision matrix: which compression technique to choose

ContextRecommendationWhy
Limited memory but quality critical (smartphone, edge inference)Pure distillation or pruning + distillationRecovers 85-97% of teacher with ×4 to ×14 reduction.
Latency critical, quality secondary (autocomplete, fast classif)Quantization alone (INT8 or INT4)No retraining, immediate memory/compute gain.
Oversized but trained architecture (under-utilized 70B model)Pruning + distillationPruning removes excess, distillation recovers precision.
Reasoning capabilities to transfer (math, code, RL)Reasoning distillation (Orca, R1)Transfers chains of thought, not just answers.
Rare or expensive training dataSynthetic data distillation (Phi-4 style)Teacher generates examples, controlled quality.

Self-distillation: learning from oneself

A counterintuitive variant: the teacher and the student are the same model, or two close versions. A model conditioned on the correct answer guides a version without this conditioning. Self-distillation via reinforcement (Self-Distilled RL, 2025) allows the model to discover reasoning patterns via its own outputs.

Surprising empirical result: self-distillation can surface capabilities absent from the initial model. The mechanism is not yet well understood theoretically, but the results suggest that distillation does not merely compress — it can also reveal.

Limits and open questions

Distillation is not magic. Several limits are documented.

Emergent capabilities resist transfer poorly. A student that is too small cannot support capabilities that require a certain architectural size to express — regardless of the quality of the supervision signal.

Teacher dependency is a risk. If the teacher produces incorrect reasoning, the student learns it. This bias propagates silently, without a standardized detection mechanism.

Black-box distillation — access only to outputs via API, without internal layers — is the dominant form in practice (access to the teacher’s weights is not always available). But it loses information from intermediate representations. Recent work (Generative Adversarial Distillation, 2025) attempts to bridge this gap, at a high computational cost.

Finally, for very small models (fewer than 3B parameters), the bottleneck is no longer supervision but the architectural capacity of the student. Beyond a certain compression threshold, improving the teacher no longer makes much difference.


Key takeaways

  • Distillation transfers knowledge from a large model (teacher) to a small one (student) via complete probability distributions — not just labels.
  • Three major families exist: logit distillation (outputs), intermediate representation distillation (hidden layers), and synthetic data distillation.
  • Reasoning distillation (Orca, DeepSeek-R1) transfers qualitative capabilities, not just competence — a conceptual leap.
  • Distillation, pruning, and quantization are complementary. The optimal sequence is pruning → distillation → quantization. Distillation is the only technique capable of recovering performance after compression.
  • Main limits: emergent capabilities resist transfer, and a student that is too small cannot support what the teacher transmits.