What's inside
The architecture without a single equation — from what one parameter is to why the second token is cheaper than the first.
13 concepts, about 78 minutes of reading at roughly six minutes each. 3 of them are free to read with no account; the rest need a paid plan.
What is in this track
- Neural network — Layers of multiply, add and bend — repeated until the numbers mean something. (free)
- Parameters and weights — A model is a very long list of decimals, and the list is all there is. (free)
- What training actually does — Guess, measure the wrongness, nudge every parameter. Repeat a few trillion times.
- The transformer — The 2017 design underneath every chatbot you have used, and why it won.
- Attention — Every word decides which other words are worth reading.
- Self-attention — Distance costs nothing. Length costs a lot — squared.
- Multi-head attention — The same mechanism run dozens of times over, each reading the sentence differently.
- Position and RoPE — Attention sees a bag of words. Order has to be added back on purpose.
- Encoder-decoder vs decoder-only — Chat models are half of the original transformer, and that is why they generalise.
- The KV cache — The first token is a whole-prompt problem. Every token after it is a one-token problem.
- Mixture of experts — Huge total parameters, small active ones — you pay memory for the size and compute for the slice.
- State space models and Mamba — A live attempt to escape quadratic attention — hybrids ship, pure ones have not taken over.
- How image models work — Image models do not draw. They start with pure noise and repeatedly remove the parts that do not look like your prompt. (free)
Before this: Memory and truth
After this: How models are made
Every track · Pricing · Claims we checked and could not stand behind