One piece at a time
The model writes a small piece of text, reads its own output, and writes one more. Hundreds of times.
Part of the What is this thing? track on lAItest.
The model does not write a sentence. It writes one piece, reads what it just wrote, and writes one more.
Then it does that a few hundred more times.
The loop
Give it "The capital of France is". It scores every possible continuation. " Paris" scores very high. It picks that, glues it on, and now the text is "The capital of France is Paris". It runs again on the longer text. And again. It stops when it produces a special piece that means stop. This is called autoregressive generation: its own output becomes its next input.
A common misconception
Commonly believed: The model works out what it is going to say, then says it.
Actually: It commits to each piece before knowing where the sentence lands. There is no plan being executed. This is why a model can open a confident sentence and finish it badly — and why letting it work a problem out in writing changes the answer. The written working is genuinely how it gets there, not a summary of thinking that already happened somewhere else.
Try it
Pick a continuation, then another. Each choice decides what can come next, and you never get to take one back. This step is an interactive widget; open the lesson to use it.
Why can a model start a sentence well and end it wrong?
Answer: It chose each piece in order and cannot take an early one back. Each piece is chosen from everything before it. Once "The answer is" has been written, the model has to continue from there. Generation runs one way, with no undo, so an early wrong turn has to be carried to the end of the sentence.
In one sentence
Every answer you have ever seen from one of these things was built one small piece at a time, left to right, with no going back.