How models are made

Pretraining, RLHF, LoRA, Chinchilla. The whole path from a pile of text to a model you can call, with the hand-waving removed.

13 concepts, about 78 minutes of reading at roughly six minutes each. 3 of them are free to read with no account; the rest need a paid plan.

What is in this track

  1. Pretraining — Before a model can follow an instruction it spends months doing one thing: guessing the next token. (free)
  2. Base models and instruct models — The model that finishes pretraining is not the model you talk to. It is the raw material. (free)
  3. Supervised fine-tuning — Show the model a few thousand examples of a request answered well, and answering well becomes its default. (free)
  4. RLHF and the reward model — People are bad at writing the perfect answer and good at picking the better of two. RLHF is built on that gap.
  5. DPO, RLAIF and Constitutional AI — Preference training has two expensive parts, a second network and a human. Each one has been removed.
  6. GRPO and verifiable rewards — When a program can check the answer, training no longer needs a human rater or a reward model.
  7. LoRA, QLoRA and PEFT — Freeze the model, train a small patch beside it, and get most of the benefit of fine-tuning for a sliver of the cost.
  8. Catastrophic forgetting — Train a model hard on your data and it can quietly lose abilities nobody thought to test.
  9. Fine-tune or prompt? — Most problems people bring to fine-tuning turn out to be prompt, retrieval or model-choice problems in disguise.
  10. Distillation — Train a small model on a big one and it inherits a surprising amount of the big one behaviour, at a fraction of the price.
  11. Synthetic data and model collapse — Training on model output is now standard practice, and whether it poisons the well turns on one detail people skip.
  12. Scaling laws and Chinchilla — Model quality moves predictably with compute, and the field spent years splitting that compute the wrong way.
  13. Emergent abilities — Some skills appear to switch on suddenly at scale. Whether that is real or an artefact of the metric is still argued.

Before this: What's inside

After this: Cheap and fast

Every track · Pricing · Claims we checked and could not stand behind