Pretraining
Before a model can follow an instruction it spends months doing one thing: guessing the next token.
Part of the How models are made track on lAItest.
Nobody taught the model facts. It only ever practised one exercise: guess what comes next.
Everything else is built on top of that.
Pretraining is one exercise, repeated at absurd scale.
Take an enormous pile of text. Hide the next token. Ask the model to predict it. Compare the guess against the real token, then nudge every weight slightly toward a better guess. Repeat until the compute budget runs out. There is no lesson plan, no labels, no questions paired with answers. Grammar, arithmetic, trivia and code style all fall out of the same drill, because all of them help you predict the next token.
Try it
The thing being predicted is a token, not a word. Type a sentence and watch where the model actually cuts it. This step is an interactive widget; open the lesson to use it.
A common misconception
Commonly believed: The model read the internet and stored it, so training is an expensive copy-paste into a giant lookup table.
Actually: The weights are far smaller than the text they were trained on, so storage was never an option. The model is forced to compress: keep the patterns that help predict many passages, discard the rest. That compression is why it generalises to sentences nobody ever wrote, and it is also why it can produce something confident that was never in the data.
What comes out of pretraining is not a chatbot.
It is a document continuer. Give it "Dear Sir," and it writes a letter. Ask it a question and it may answer with three more questions, because a page containing one question usually contains several. Turning that into something that answers you is a separate job, done later, with different data.
In one sentence
Pretraining does not teach a model to be helpful. It teaches a model to predict, and helpfulness has to be added afterwards.