What training actually does

Guess, measure the wrongness, nudge every parameter. Repeat a few trillion times.

Part of the What's inside track on lAItest.

Training a model is one move, repeated a few trillion times.

The move is: guess, measure how wrong, nudge.

Step one: guess the next word.

Take a chunk of real text, hide what comes next, and let the network guess. It does not produce one word — it produces a score for every word in its vocabulary at once. At the start, when the parameters are random, those scores are noise.

Step two: squash the wrongness into a single number.

That number is the loss. It is high when the network gave the true next word a low score, and low when it gave it a high score. Everything else in training exists to push that one number down. Choosing what the loss measures is therefore the single biggest decision anyone makes about a model.

Step three: nudge every parameter a little.

Backpropagation runs backwards through the layers and works out, for each parameter, which direction would have made the loss smaller. Gradient descent then moves each one a small step that way. A single step barely helps. Repeated across enormous amounts of text, the network stops guessing and starts predicting.

A common misconception

Commonly believed: Someone has to label the training data, the way you label photos to teach an image classifier.

Actually: Nobody labels anything. The text supplies its own answer key, because the next word is always sitting right there in the document. That is why the method scales: any text ever written is training data without a human writing a single question.

What is the "loss" during training?

Answer: A single number saying how wrong one guess was. Loss compresses a whole guess into one number so that backpropagation has something to differentiate. Nothing in the model knows what a fact is — it only knows which way to move to make that number smaller.

In one sentence

Training is not teaching. It is billions of tiny corrections to a score for being wrong.