Streaming and time to first token
Streaming does not make generation faster. It makes the waiting shorter, which is a different product.
Part of the Turning the dials track on lAItest.
The answer took four seconds either way. Streaming only stopped you staring at nothing for four seconds.
Streaming sends each token as it is produced instead of holding them all back.
Generation is sequential regardless: one token, then the next, then the next. Without streaming the server keeps everything until the model finishes and sends a single block. With streaming it forwards each piece over an open connection as it arrives. Total time barely changes. What changes is when the first word reaches a human.
Time to first token is the number people actually feel.
TTFT measures the gap between sending a request and receiving the first token back. It covers the network trip, any queueing, and the work of reading your entire prompt before a single output token exists. That last part is why a long prompt raises TTFT even when the answer is short. After the first token, the rest tend to arrive at a fairly steady rate.
Try it
Turn streaming on and watch where the first sound lands. The work does not shrink; the wait does. This step is an interactive widget; open the lesson to use it.
You paste a long document into the prompt. The answer stays two sentences. What gets slower?
Answer: Time to first token. The model must read the whole prompt before it can emit anything, so prompt length lands almost entirely on time to first token. Once output has started, the per-token rate is largely unaffected by how much you sent in.
In one sentence
Streaming does not speed up generation. It shortens the wait, and shortening the wait is what users notice.