Multi-head attention

The same mechanism run dozens of times over, each reading the sentence differently.

Part of the What's inside track on lAItest.

One set of attention scores cannot track grammar and meaning at the same time.

So a layer runs dozens of them side by side.

A head is one independent set of scores.

Each word vector is split into slices. Every head gets its own query, key and value projections and works out its own weights over the passage. The results are stitched back together and mixed. Heads are never told what to specialise in — they land on different jobs because different jobs happen to reduce the loss.

Try it

Switch heads on one sentence. Grammar, reference and nearby-words are three different readings of identical text. This step is an interactive widget; open the lesson to use it.

A common misconception

Commonly believed: Researchers assign each head a job: this one does grammar, that one does pronouns.

Actually: Nothing is assigned. Interpretability work finds, after the fact, heads that behave reliably like grammar heads or reference heads — and plenty of heads with no clean story at all. The specialisation is something discovered in a trained model, not something designed into it.

In one sentence

Multi-head attention is one mechanism run many times in parallel, each copy reading the same sentence for something different.