Position and RoPE
Attention sees a bag of words. Order has to be added back on purpose.
Part of the What's inside track on lAItest.
To bare attention, "dog bites man" and "man bites dog" are the same input.
Word order is not something the mechanism knows.
Scoring every word against every word never mentions where they are.
The comparison is between two vectors. It carries no index, no sense of before or after. Shuffle the passage and every score comes out identical. So position has to be written into the numbers themselves, before attention ever runs.
RoPE encodes position as a rotation.
Rotary position embedding (arXiv:2104.09864) rotates each word's query and key vectors by an angle proportional to that word's position. Because a score depends on the angle between two vectors, the score ends up depending on how far apart the words are rather than where they sit absolutely. It is the standard choice across current open architectures.
What breaks in a transformer given no positional information at all?
Answer: It cannot tell a sentence apart from the same words shuffled. Attention scores are computed pairwise with no notion of index, so a shuffled passage scores identically. Position lives in the vectors — which is also why pushing a model past the length it was trained on is a position problem before it is anything else.
In one sentence
Attention sees an unordered bag of words; positional encoding is what turns it back into a sentence.