What training actually does
Guess, measure the wrongness, nudge every parameter. Repeat a few trillion times.
Part of the What's inside track on lAItest.
Training a model is one move, repeated a few trillion times.
The move is: guess, measure how wrong, nudge.
Step one: guess the next word.
Take a chunk of real text, hide what comes next, and let the network guess. It does not produce one word — it produces a score for every word in its vocabulary at once. At the start, when the parameters are random, those scores are noise.
Step two: squash the wrongness into a single number.
That number is the loss. It is high when the network gave the true next word a low score, and low when it gave it a high score. Everything else in training exists to push that one number down. Choosing what the loss measures is therefore the single biggest decision anyone makes about a model.
Step three: nudge every parameter a little.
Backpropagation runs backwards through the layers and works out, for each parameter, which direction would have made the loss smaller. Gradient descent then moves each one a small step that way. A single step barely helps. Repeated across enormous amounts of text, the network stops guessing and starts predicting.
A common misconception
Commonly believed: Someone has to label the training data, the way you label photos to teach an image classifier.
Actually: Nobody labels anything. The text supplies its own answer key, because the next word is always sitting right there in the document. That is why the method scales: any text ever written is training data without a human writing a single question.
What is the "loss" during training?
Answer: A single number saying how wrong one guess was. Loss compresses a whole guess into one number so that backpropagation has something to differentiate. Nothing in the model knows what a fact is — it only knows which way to move to make that number smaller.
In one sentence
Training is not teaching. It is billions of tiny corrections to a score for being wrong.