Turning the dials
Logits, softmax, temperature, top-p, streaming and stop conditions — the whole control panel, without the maths anxiety.
10 concepts, about 60 minutes of reading at roughly six minutes each. Free to read, with no account.
What is in this track
- Logits — A model never picks a word. It scores every word it knows, and something else picks. (free)
- Softmax — The step that turns a pile of raw scores into percentages that add up to 100. (free)
- Sampling and decoding — Same prompt, same model, different answer. The model did not change its mind; the draw did. (free)
- Temperature — One number decides how far the sampler is allowed to stray from the model’s favourite token. (free)
- Top-p and top-k — Temperature reshapes the whole list. These two just delete the bottom of it. (free)
- Determinism — Temperature 0 removes the dice. It does not make the system reproducible. (free)
- Stop sequences and max_tokens — A model has no sense of how long your answer should be. Two settings decide when it ends. (free)
- Streaming and time to first token — Streaming does not make generation faster. It makes the waiting shorter, which is a different product. (free)
- Throughput versus latency — A serving system can be fast for one person or fast for everyone, and the two fight each other. (free)
- Perplexity — One number for how surprised a model is by a piece of text. It travels much less well than people assume. (free)
Before this: What is this thing?
After this: Getting good answers
Every track · Pricing · Claims we checked and could not stand behind