Cheap and fast

Where the money actually goes: memory, batches, caches, tiers and routing. How to cut a bill by an order of magnitude without changing the product.

12 concepts, about 72 minutes of reading at roughly six minutes each. 1 of them is free to read with no account; the rest need a paid plan.

What is in this track

  1. What you actually rent — Serving a model is a memory problem wearing a compute costume. The weights have to sit somewhere the chip can reach. (free)
  2. Precision and quantization — Every weight costs bits. Spend fewer bits per weight and the model shrinks, speeds up, and gets slightly worse in ways you have to measure.
  3. Batching and the KV cache — Your token price is a group rate. The size of the group is capped by how much conversation the memory can hold.
  4. Speculative decoding — A small model guesses the next few tokens and the big model grades them in one pass. Accepted guesses are exactly what the big model would have written.
  5. Prompt caching — The unchanging front of your prompt can be billed at a fraction of the normal rate — as long as it genuinely never changes.
  6. The Batch API — A discount for giving up latency, nothing else. Same model, same quality, results collected later.
  7. The spread — Models of comparable capability differ in price by more than an order of magnitude. Both prices are real, and neither is a quality signal.
  8. Tiers and surcharges — A quoted rate assumes a short request on the standard tier. Tier and length are separate multipliers, and they compound.
  9. Model routing — Most requests in a real product are easy. Sending all of them to your best model is the most expensive habit in the industry.
  10. Small and on-device — Small models win on latency, privacy, offline operation and licence — not on difficulty. Self-hosting to save money often loses to the hosted floor.
  11. Limits, retries and refusals — Rate limits come in two shapes, retries need idempotency to be safe, and a refusal can arrive as a perfectly successful response.
  12. Deprecation and versioning — Pin the model id, watch the retirement calendar, and treat a version bump as a change that needs its own evals.

Before this: How models are made

After this: Agents

Every track · Pricing · Claims we checked and could not stand behind