Distillation

Train a small model on a big one and it inherits a surprising amount of the big one behaviour, at a fraction of the price.

Part of the How models are made track on lAItest.

A small model can learn more from a large model answering a question than from the original human text.

Same words, more signal.

The teacher does not only give an answer. It gives a whole distribution.

Ask a big model to continue a sentence and it does not merely pick one token; it assigns a probability to every token it might have picked. Those runner-up probabilities carry information the raw text never did, that two words were nearly interchangeable here, that a third was almost as good. Training the small model to match the whole distribution rather than just the top choice is the original distillation idea. Paper: arXiv 1503.02531.

In practice, most of what gets called distillation is simpler than that.

Generate a large pile of outputs from a strong model and fine-tune a smaller one on them as if a person had written them. It is supervised fine-tuning where the demonstrator is a machine. It works well enough to be everywhere, and because the teacher produces far more examples than any annotation team could, the ceiling becomes the quality of the teacher.

A common misconception

Commonly believed: A distilled model is a compressed copy of the teacher, so it should behave the same, only faster.

Actually: It copies behaviour on the kinds of prompt it was distilled from, and nowhere else. Push it off that distribution and the imitation thins out quickly, because the student learned to sound like the teacher rather than to reason like it. Distillation is also where terms of service bite: several vendors explicitly forbid using their outputs to train a competing model.

Try it

Sort the live index by input price. The distance between the cheapest usable model and the most expensive is the whole reason distillation exists. This step is an interactive widget; open the lesson to use it.

What a small model costs against a frontier one

Read from a live model index at page-render time, each figure linked to the vendor page it came from.
WhatValueProvenance
Claude Opus 5, input$5 per 1Msource, verified .
Claude Haiku 4.5, input$1 per 1Msource, verified .
DeepSeek Flash, input$0.30 per 1Msource, verified .
Cheapest frontier input price$0.75 per 1Msource, verified .

In one sentence

Distillation moves behaviour from a big expensive model into a small cheap one, and behaviour is not the same thing as capability.