Batching and the KV cache
Your token price is a group rate. The size of the group is capped by how much conversation the memory can hold.
Part of the Cheap and fast track on lAItest.
One user on a GPU is a wasted GPU. The price you pay assumes you are sharing.
Batching spreads the fetch across everybody
If the weights must be read anyway, read them once and push many requests through together. Every request in the batch shares a cost that one request alone would have paid by itself. That is most of the reason a million tokens costs a few dollars instead of a few hundred. It is also why throughput and latency pull against each other: a bigger batch serves more people per second and makes each of them wait slightly longer.
Try it
Fill the window and watch what has to be held in memory for every token already in it. This step is an interactive widget; open the lesson to use it.
The KV cache is what caps the batch
Attention lets each new token look back at every earlier one. To avoid redoing that work on every step, the server keeps a small record for each token in the conversation. That is the KV cache. It sits in the same scarce memory as the weights and it grows with every token. A long conversation does not only cost more tokens. It takes seats away from other users.
PagedAttention gave the seats back
Early servers reserved one continuous block of memory per request, sized for the longest answer that request might produce. Most of that reservation was never used. vLLM introduced PagedAttention (arXiv:2309.06180), which stores the cache in small fixed-size blocks handed out as they are needed, the way an operating system hands out pages of memory. More requests fit on the same card, so each one costs less.
In one sentence
Cheap tokens are a side effect of crowding. Anything that makes a request hog memory makes the crowd smaller.