What you actually rent
Serving a model is a memory problem wearing a compute costume. The weights have to sit somewhere the chip can reach.
Part of the Cheap and fast track on lAItest.
A model is not thinking too hard. It is running out of memory.
The real constraint is where the weights sit.
Weights have to live somewhere
A model is a large pile of numbers called weights. To answer you, a chip has to read them. Reading them from the memory attached to that chip is fast. Reading them from anywhere else is so slow it is not worth doing. So the first question about running any model is not how smart it is. It is whether its weights fit in the memory you rented.
Mixture of experts changes what runs, not what you store
Kimi K3 holds about 2.8 trillion weights and uses roughly 104 billion of them for any single token. That saves arithmetic. It does not save storage. When Moonshot published the weights in July 2026, the download was about 1.56 terabytes spread across 96 shards. All of it still has to be somewhere a chip can reach quickly.
A common misconception
Commonly believed: Running a model is expensive because generating text takes an enormous amount of arithmetic.
Actually: Chips are fast at arithmetic and comparatively slow at fetching. Producing one token means reading the weights it needs, and then the next token means reading them again. The fetching is the bottleneck. That is why serving one user at a time is the most wasteful possible way to run a model.
In one sentence
You are not renting intelligence by the hour. You are renting memory, and the bill is set by how busy that memory is kept.