Encoder-decoder vs decoder-only

Chat models are half of the original transformer, and that is why they generalise.

Part of the What's inside track on lAItest.

The original transformer was a translation machine. Chat models threw half of it away.

The half they kept is the half that can write.

Two halves, two jobs.

The encoder reads the whole input at once, every word free to look left and right, and produces a representation of it. The decoder writes the output one token at a time. It may look at everything the encoder produced, but only at output words it has already written. That mask is what makes generation possible at all: you cannot read a word you have not written yet.

Today's chat models are decoder only.

Drop the encoder entirely and treat the prompt as the opening of a document the model is finishing. Prompt and reply live in one sequence under one mechanism. It is simpler to build, it trains on any text at all rather than on paired examples, and it is the shape of every model you have held a conversation with.

A common misconception

Commonly believed: Decoder-only models won because they are more powerful than encoder-decoder ones.

Actually: They won because they are more general. Encoder-decoder models remain excellent at fixed input-to-output jobs such as translation. The decoder-only shape simply has no task boundary to define, so a single model absorbs every task that can be phrased as text continuing text.

In one sentence

A chat model is one long document it keeps finishing. Your prompt is just the part already written.