Supervised fine-tuning

Show the model a few thousand examples of a request answered well, and answering well becomes its default.

Part of the How models are made track on lAItest.

The cheapest way to turn a text continuer into an assistant is to show it a few thousand examples of one.

No reinforcement learning required.

Supervised fine-tuning is pretraining with curated pages.

The mechanics do not change: hide the next token, predict it, nudge the weights. Only the data changes. Instead of scraped documents the model sees pairs, a request and a response somebody was willing to sign their name to. Because responses now reliably follow requests, answering becomes the high-probability continuation. Same loss function, pointed at a much smaller and much more expensive pile of text.

Instruction tuning is supervised fine-tuning with variety as the point.

Here the examples deliberately span many task types: summarise this, translate this, fix this code, decline this. The aim is not to memorise those tasks. It is to teach the general shape of "a person asked for something, here is a direct attempt at it", so the model handles requests it never saw during tuning.

A common misconception

Commonly believed: Fine-tuning data has to be enormous, the way pretraining data is enormous.

Actually: It is the opposite trade. Pretraining wants scale and tolerates junk. Fine-tuning wants a small, clean, internally consistent set, because every example is a direct demonstration of the behaviour you are asking for. A few thousand carefully written responses can change how a model talks. A few thousand sloppy ones will teach it to be sloppy, precisely and durably.

Where demonstrations run out.

A demonstration shows the model one good answer. It never says why the answer was good, and it never says what would have been slightly better or much worse. Whoever writes the example also has to be capable of producing the answer themselves. When you want the model to exceed the person supervising it, or when quality is easier to judge than to write, demonstrations stop being enough.

In one sentence

Supervised fine-tuning teaches by demonstration: here is a request, here is a response worth copying.