Model Distillation Explained
How a small model inherits the behaviour of a much larger one, what gets lost along the way, and why distillation is the reason cheap models got good so fast.
Distillation is the process of training a small model to reproduce the behaviour of a large one. The large model is the teacher, the small one is the student, and the training signal comes from the teacher rather than from raw text.
It is the main reason a 7B model in 2026 can do things that needed a frontier model two years earlier, and the main reason the relationship between parameter count and capability keeps getting weaker.
Why it works at all
The intuition is that a trained model contains far more information than its final answers reveal.
Ask a model to classify a piece of text and it does not simply pick a label — internally it produces a probability distribution over every possible next token. That distribution encodes which alternatives were plausible and how close the call was. Training a student on the full distribution transfers much more than training it on the winning label alone.
This is the original insight behind distillation, and it is sometimes described as learning from the teacher uncertainty. A ground-truth label says "this is a cat". A teacher distribution says "this is a cat, but it is somewhat dog-like and not at all car-like", which is a richer lesson from the same example.
The three ways it is actually done
Logit distillation. The student is trained to match the teacher output distribution directly, minimising the divergence between them. This is the classical form. It requires access to the teacher raw outputs, so in practice it is available only to whoever owns the teacher.
Sequence-level or synthetic-data distillation. The teacher generates a large corpus of outputs — answers, reasoning traces, tool-use trajectories — and the student is fine-tuned on those as if they were human-written training data. This needs only sampled text, so it works against any model you can call.
This is the dominant form today, and the publicly documented example is instructive: the DeepSeek-R1 release included a family of dense models fine-tuned on reasoning data generated by the larger R1 model, built on Qwen and Llama base checkpoints at 1.5B, 7B, 8B, 14B, 32B and 70B, with the reported results showing distilled models outperforming considerably larger non-distilled ones on reasoning benchmarks.
Intermediate-representation distillation. The student is trained to match the teacher internal activations layer by layer, not just its outputs. It transfers more but requires architectural compatibility and full white-box access.
What gets lost
Distillation compresses. Compression is lossy, and the losses are systematic rather than random.
- The tail. Students learn the common cases well because that is where the teacher data concentrates. Rare domains, unusual languages and edge-case reasoning degrade first, and they degrade quietly.
- Robustness under distribution shift. A student matched to the teacher on the distillation set can diverge sharply on inputs unlike anything in that set.
- Calibration. Distilled models are frequently more confidently wrong than their teacher. They inherited the shape of the answer without the underlying uncertainty.
- Long-horizon consistency. The gap is smallest on single-step tasks and largest on multi-step agent work, where small per-step errors compound.
The consequence for anyone selecting a model: a distilled model that matches its teacher on a benchmark will not necessarily match it on your workload, because the benchmark is probably closer to the distillation distribution than your workload is.
Distillation versus the things it is confused with
Quantisation reduces the numeric precision of an existing model — same parameters, fewer bits each. The model is unchanged in structure; it is stored more coarsely. Distillation produces a genuinely different, smaller model.
Pruning removes parameters judged unimportant from a trained model. Also structural, also not distillation, and the two are often combined.
Fine-tuning adapts a model to a task using task data. Distillation is fine-tuning where the training data came from another model. The mechanics overlap; the purpose does not.
All four are used together in practice. A production small model is frequently distilled from a large teacher, then quantised for serving.
Why this matters commercially
Two consequences shape the market you are buying in.
First, the cost of capability falls faster than the cost of compute. Each generation of frontier models becomes a teacher for the next generation of small ones, so the price of a given capability level drops on a much steeper curve than hardware improvement alone would produce. A model you are paying frontier rates for today has a cheap approximate equivalent sooner than you expect.
Second, model provider terms of service almost universally prohibit using their outputs to train competing models. Whether that is enforceable is a live legal question and not one to resolve in your product roadmap. If you are distilling from a commercial API, read the terms; if you are distilling from open-weight models, read the licence, which is frequently permissive — several current open-weight releases carry MIT licences.
Distilling your own
For a narrow, high-volume task this is one of the highest-return optimisations available. The recipe is unglamorous:
- Collect real production inputs — a few thousand, sampled rather than curated.
- Generate outputs with the strongest model you can justify, including reasoning traces if the task benefits.
- Filter aggressively. Bad teacher outputs become permanent student behaviour, and filtering is where most of the quality comes from.
- Fine-tune a small open-weight base model on the filtered set.
- Evaluate against held-out real inputs, not against the teacher. Matching the teacher is not the goal; doing the task is.
The economics only work when volume is high and the task is stable. For varied, evolving work, calling a general model is cheaper than maintaining a bespoke one — and if the reason you are considering distillation is per-token cost on exploratory work, flat-rate access solves that problem without a training pipeline to own.
Common questions
Is a distilled model as good as its teacher?
On the tasks it was distilled for, often close. On rare cases, unusual inputs and long multi-step work, noticeably worse. The losses are systematic rather than random, which is why evaluation on your own inputs matters more than benchmark parity.
What is the difference between distillation and quantisation?
Quantisation stores the same parameters at lower numeric precision, leaving the model structurally unchanged. Distillation trains a genuinely smaller model to imitate a larger one. They are complementary and frequently applied together.
Can I legally distil from a commercial model API?
Most commercial provider terms prohibit using their outputs to train competing models, so check the terms before building anything on it. Distilling from open-weight models is governed by the model licence instead, and several current releases are permissively licensed.