Training, then running

Training is the enormous one-off process that set the numbers. Inference is what happens each time you press send.

Part of the What is this thing? track on lAItest.

The expensive part happened months before you opened the app, and for that model it will never happen again.

What you are using is the leftovers. The leftovers are the point.

Training, very roughly

Take a mountain of text. Hide the next chunk. Make the model guess it. Compare the guess with what the text actually said. Nudge the numbers a tiny amount in the direction that would have made the guess better. Repeat that an unimaginable number of times over an unimaginable amount of text. Nobody types in any facts. Facts are a side effect of getting very good at guessing text that contains them.

Then it is taught to be useful

A model trained only that way is a text continuer: ask it a question and it might reply with more questions, because documents full of questions are a thing that exists. So a second, much smaller stage follows, where people and other models show it what a good response to a request looks like and rate its attempts. That stage is what turns a continuer into something that answers you. A later track takes it apart properly.

Inference is the part you pay for

Running a finished model on your input is called inference. The numbers do not change. Nothing is written down. It costs some electricity and a slice of a very expensive chip, every single time — which is why usage is billed per token rather than sold as a one-off, and why the same model can be offered fast and dear or slow and cheap.

What changes inside the model while you are chatting with it?

Answer: Nothing — the numbers are frozen. Training and inference are separate events. Once training ends the numbers are fixed, and every conversation in the world runs against the identical model. Anything that feels like personalisation is text being added to the prompt, not numbers being changed.

In one sentence

Training built the thing once, at enormous cost. Inference is running it — and running it is the only thing that ever happens to you.