State space models and Mamba
A live attempt to escape quadratic attention — hybrids ship, pure ones have not taken over.
Part of the What's inside track on lAItest.
Attention costs grow with the square of the input. Somebody was always going to attack that.
State space models are the serious attempt.
Carry a fixed-size summary instead of the whole past.
A state space model keeps a state of fixed size and updates it as it moves through the sequence. Cost per token stays constant however long the passage gets, and there is no cache growing without bound. Mamba (arXiv:2312.00752) made that state selective: what gets written into it depends on the token being read.
The trade is exact recall against cost.
Attention keeps every earlier token available for exact lookup and pays for it in memory and arithmetic. A fixed-size state cannot keep everything, so something must be forgotten — and what gets forgotten is decided by learned parameters, before your question arrives, rather than by the question itself.
A common misconception
Commonly believed: State space models replaced the transformer; attention is on the way out.
Actually: Nothing you talk to today is a pure state space model. Where the idea ships in production it ships as a hybrid — a few attention layers among many state space layers — because a fixed state is the wrong tool for quoting an exact line out of a long document and the right tool for almost everything around it. This is an open contest, not a finished one.
In one sentence
Attention remembers everything expensively; state space models remember cheaply and imperfectly, and the field has not finished arguing about the exchange rate.