The transformer

The 2017 design underneath every chatbot you have used, and why it won.

Part of the What's inside track on lAItest.

One 2017 paper sits underneath every chat model you have ever used.

Attention Is All You Need — arXiv:1706.03762.

What it removed mattered more than what it added.

The language models before it read a sentence one word at a time, carrying a running summary forward. That made training strictly sequential: word two could not be processed until word one finished. The transformer threw the running summary away. Every position is processed at the same moment, and each one looks directly at all the others.

Reading everything at once is why the models got large.

A design that can process a whole passage in parallel can use thousands of chips at the same time. The sequential design could not. Scale stopped being a research problem and became a spending problem — and that shift, more than any accuracy result, produced the models we have now.

A common misconception

Commonly believed: The transformer won because it was simply a better model of language.

Actually: It was better on its benchmark, but the durable reason is that it trains well on hardware you can buy by the rack. Architectures that suit parallel machines get scaled until they look brilliant. Architectures that do not, quietly never get the chance.

What did the transformer stop doing that earlier language models did?

Answer: Processing a passage strictly one word at a time. It still predicts the next word, it is still a neural network, and it still stores parameters. What it dropped is the forced left-to-right pass, which is precisely what let one training run occupy a whole datacentre.

In one sentence

The transformer won because it fits the hardware, not only because it fits the language.