How models are made
Pretraining, RLHF, LoRA, Chinchilla. The whole path from a pile of text to a model you can call, with the hand-waving removed.
13 concepts, about 78 minutes of reading at roughly six minutes each. 3 of them are free to read with no account; the rest need a paid plan.
What is in this track
- Pretraining — Before a model can follow an instruction it spends months doing one thing: guessing the next token. (free)
- Base models and instruct models — The model that finishes pretraining is not the model you talk to. It is the raw material. (free)
- Supervised fine-tuning — Show the model a few thousand examples of a request answered well, and answering well becomes its default. (free)
- RLHF and the reward model — People are bad at writing the perfect answer and good at picking the better of two. RLHF is built on that gap.
- DPO, RLAIF and Constitutional AI — Preference training has two expensive parts, a second network and a human. Each one has been removed.
- GRPO and verifiable rewards — When a program can check the answer, training no longer needs a human rater or a reward model.
- LoRA, QLoRA and PEFT — Freeze the model, train a small patch beside it, and get most of the benefit of fine-tuning for a sliver of the cost.
- Catastrophic forgetting — Train a model hard on your data and it can quietly lose abilities nobody thought to test.
- Fine-tune or prompt? — Most problems people bring to fine-tuning turn out to be prompt, retrieval or model-choice problems in disguise.
- Distillation — Train a small model on a big one and it inherits a surprising amount of the big one behaviour, at a fraction of the price.
- Synthetic data and model collapse — Training on model output is now standard practice, and whether it poisons the well turns on one detail people skip.
- Scaling laws and Chinchilla — Model quality moves predictably with compute, and the field spent years splitting that compute the wrong way.
- Emergent abilities — Some skills appear to switch on suddenly at scale. Whether that is real or an artefact of the metric is still argued.
Before this: What's inside
After this: Cheap and fast
Every track · Pricing · Claims we checked and could not stand behind