How image models work
Image models do not draw. They start with pure noise and repeatedly remove the parts that do not look like your prompt.
Part of the What's inside track on lAItest.
The picture was already in the noise. The model only had to take away everything that was not it.
Nothing here predicts a next token.
Learn to destroy, then run it backwards.
Training shows the model real images with a measured amount of random noise added, over and over, until an image is indistinguishable from static. The model learns one narrow skill: given a noisy image, estimate the noise. That is all. Nothing about it can draw.
Generation is that skill, applied repeatedly.
Start with pure static. Ask the model what noise it sees. Subtract a fraction of it. Ask again. After a few dozen rounds an image is left behind — one that was never in the training set, because at every step the guess was steered by your prompt. Denoising diffusion was set out in arXiv:2006.11239.
Try it
Drag from pure noise to a finished picture, and watch structure arrive before detail does. This step is an interactive widget; open the lesson to use it.
A common misconception
Commonly believed: It is collaging pieces of pictures it has seen, or looking up something close and editing it.
Actually: There is no image storage and no lookup. What was learned is a function from noisy pixels to a noise estimate, which is why the same prompt with a different starting static gives a genuinely different picture — and why more denoising steps buy detail at the cost of time.
Why does the same prompt produce a different image each time?
Answer: The starting noise is different, and everything downstream follows from it. The prompt only steers the denoising. The starting static is the seed of the whole picture, which is why fixing that seed makes the output reproducible.
In one sentence
A language model adds one piece at a time; an image model removes noise many times over. Same idea of learned prediction, opposite direction of travel.