Self-attention

Distance costs nothing. Length costs a lot — squared.

Part of the What's inside track on lAItest.

Every word in your prompt reads every other word. All of them. In every layer.

That is where the bill comes from.

Query, key, value.

Each word turns its own numbers into three vectors. The query is what this word is looking for. The key is what this word offers. The value is what it hands over if chosen. Score every query against every key, turn the scores into weights, and mix the values in those proportions. "Self" means all three come from the same passage.

Distance is free. Length is not.

A word six positions back and a word six hundred positions back are both a single hop away. That is the real advantage over reading in order. But comparing every word to every word means the work grows with the square of the length: double the passage and you quadruple the attention arithmetic.

Try it

Focus "cooked". Its subject sits six words back, and the model reaches it in one step rather than six. This step is an interactive widget; open the lesson to use it.

Why does doubling the length of a prompt more than double the attention work?

Answer: Every word is compared against every word, so comparisons grow with the square of the length. Layers and parameters are fixed no matter what you send. The passage is not. Roughly n by n comparisons, per head, per layer, is why long context is a real engineering cost rather than a bigger box.

In one sentence

Self-attention gives you distance for free and charges you for length, squared.