Cheap and fast
Where the money actually goes: memory, batches, caches, tiers and routing. How to cut a bill by an order of magnitude without changing the product.
12 concepts, about 72 minutes of reading at roughly six minutes each. 1 of them is free to read with no account; the rest need a paid plan.
What is in this track
- What you actually rent — Serving a model is a memory problem wearing a compute costume. The weights have to sit somewhere the chip can reach. (free)
- Precision and quantization — Every weight costs bits. Spend fewer bits per weight and the model shrinks, speeds up, and gets slightly worse in ways you have to measure.
- Batching and the KV cache — Your token price is a group rate. The size of the group is capped by how much conversation the memory can hold.
- Speculative decoding — A small model guesses the next few tokens and the big model grades them in one pass. Accepted guesses are exactly what the big model would have written.
- Prompt caching — The unchanging front of your prompt can be billed at a fraction of the normal rate — as long as it genuinely never changes.
- The Batch API — A discount for giving up latency, nothing else. Same model, same quality, results collected later.
- The spread — Models of comparable capability differ in price by more than an order of magnitude. Both prices are real, and neither is a quality signal.
- Tiers and surcharges — A quoted rate assumes a short request on the standard tier. Tier and length are separate multipliers, and they compound.
- Model routing — Most requests in a real product are easy. Sending all of them to your best model is the most expensive habit in the industry.
- Small and on-device — Small models win on latency, privacy, offline operation and licence — not on difficulty. Self-hosting to save money often loses to the hosted floor.
- Limits, retries and refusals — Rate limits come in two shapes, retries need idempotency to be safe, and a refusal can arrive as a perfectly successful response.
- Deprecation and versioning — Pin the model id, watch the retirement calendar, and treat a version bump as a change that needs its own evals.
Before this: How models are made
After this: Agents
Every track · Pricing · Claims we checked and could not stand behind