The KV cache

The first token is a whole-prompt problem. Every token after it is a one-token problem.

Part of the What's inside track on lAItest.

The five hundredth word of a reply costs the model less than the first.

It refuses to do the same work twice.

Keys and values never change once computed.

To produce the next token, the model needs a key and a value for every token before it. Those depend only on what came earlier, so the ones computed at token 400 are still exactly right at token 401. Keep them and each step computes keys and values for one new token instead of recomputing thousands.

The catch is memory, not arithmetic.

The cache holds a key and a value for every token, in every layer, in every head, for every request being served at once. It grows with the length of the conversation. On a loaded server it is the cache, not the parameters, that runs out first. Attention kernels such as FlashAttention (arXiv:2205.14135) exist largely to keep this traffic off slow memory.

How much conversation the cache may have to hold

Read from a live model index at page-render time, each figure linked to the vendor page it came from.
WhatValueProvenance
Largest frontier context window1.1Msource, verified . every token in here needs a cached key and value
Claude Opus 5 context1Msource, verified .
Claude Haiku 4.5 context200Ksource, verified .

Try it

Pour a novel into a small window, then a large one. The cache has to hold every token still visible. This step is an interactive widget; open the lesson to use it.

A common misconception

Commonly believed: Sending the same long conversation again is nearly free, because the model has already seen it.

Actually: Different request, different cache. By default a provider recomputes your entire prompt from scratch on every call. Prompt caching exists as a separate, opt-in, separately billed feature precisely because that is the default — and it is why any change near the start of a prompt is more expensive than a change near the end.

In one sentence

The KV cache is why generation gets cheaper as it goes, and why long conversations get heavier as they go.