Throughput versus latency

A serving system can be fast for one person or fast for everyone, and the two fight each other.

Part of the Turning the dials track on lAItest.

The thing that makes a model server fast for everyone is the same thing that makes it slow for you.

Latency is how long one request takes. Throughput is how many requests get served per second.

They pull against each other. Accelerators are far more efficient running many requests together in one batch, so a server waits a moment to collect requests before it runs them. That wait is pure added latency for whoever arrived first and pure gain for total capacity. Every hosted model you call sits somewhere on that trade, at a point the provider chose, not you.

Try it

Remove hops and raise speed until the first reply lands inside the window where a pause stops feeling like a pause. This step is an interactive widget; open the lesson to use it.

In a conversation, only latency is felt.

The published thresholds for voice are blunt. Under about 300 milliseconds a reply feels human. Around 500 milliseconds is the usual engineering target. Past roughly 600 milliseconds callers start pressing buttons instead of talking, and past a second and a half they hang up. Typical production turn latency still sits well above those targets, because a turn is the sum of every hop, not the fastest one.

A common misconception

Commonly believed: The vendor page says its component responds in tens of milliseconds, so replies will arrive that fast.

Actually: That figure is one component measured alone. A real turn adds the network trip, deciding the speaker has actually finished, transcription, the model’s own time to first token, speech synthesis, and buffering on the device. Measure end to end, and watch the slow tail rather than the median — users remember the bad turns.

In one sentence

Nobody experiences throughput. They experience the gap before the reply, so measure the whole turn rather than the quickest hop in it.