Linear attention and state space models replace attention with a memory of fixed size that updates as the model reads. That is much cheaper than attention and much worse at looking things up, which is why models built this way usually keep some attention layers alongside the recurrence. We tested the recurrent cells directly to find which parts of them are responsible for the shortfall. Most of it comes from a small piece most of them leave out, and from training that never showed them the situation they were tested on.
Most large language models use attention, which lets them look back at any word they have already read. That is why they are good at remembering, and it is also why they get expensive. The cost grows with the length of the text.
The alternative is a recurrent layer. Rather than letting every word look at every other word, it reads left to right and folds what it has seen into a memory of fixed size. A linear attention layer, or a state space model such as Mamba, is built from one of these. It is much cheaper than attention and much worse at looking things up. In practice that is why models built this way usually keep some attention layers as well.
The standard explanation is that the fixed size is the flaw. Once a long text has been compressed into a small state, the detail you wanted is gone. That explanation is mostly wrong. The cells are not running out of room. Two things account for nearly all of the gap. The first is a small filter over the last few words. Mamba-2 includes one, Gated DeltaNet does not, and adding it moves recall from 0.52 to 0.98 at no memory cost. The second is the training data. The cells were never shown the hard version of the task, so they never learned to handle it.
The second finding is the one worth pausing on. Nothing changed except the order and difficulty of the training examples. The architecture, the memory and the parameters were identical. A task the model scored 0.021 on, where guessing scores 0.023, became a task it solved perfectly.
Comparing these models is harder than it sounds. Mamba-2 ships with a short convolution, a filter over the last few inputs. Gated DeltaNet does not have one. RWKV-7 has a decay mechanism the others lack. Run them against each other and you learn which one wins, because three differences move at once and you cannot tell them apart.
So we rebuilt every model ourselves in plain PyTorch, gave them all exactly the same amount of memory (1,024 state elements, about 29k parameters), and then switched one part at a time. Rebuilding them also meant we could run Mamba-2 without its optimised kernel, which keeps the comparison fair.
The test itself is simple. The model reads a list of key and value pairs, then a stretch of filler text, then one of the keys. It has to produce the matching value. It is the smallest task that separates models which can look things up from models that cannot.
Put the pieces together and the best model is a gated delta rule wearing the same small filter Mamba-2 ships. At the hardest memory setting it reaches 0.983, with a worst seed of 0.917 across ten runs, and the scrambled-pairings control drops it to 0.037, which is guessing.
We also ran it against Mamba-2 using its own optimised kernel on the same GPU. At a shared default learning rate the official kernel scores 0.867 and ours 0.967. Tuned separately, both improve, to 0.955 against 1.000. The gap between them is +0.045. Most of the distance between our rebuilt version and real Mamba turned out to be the learning rate. Almost none of it was the architecture.
Ask these models to hold more pairs and they degrade gently. Everything manages 8 pairs, and the weaker models fade out toward 32. That is a real capacity curve, and it behaves the way you would expect.
Ask the same models to find 4 pairs inside a paragraph of similar-looking distractions and they stop dead. All 45 runs landed between 0.014 and 0.023, against a guessing floor of 0.019. Attention scored 1.000 at every length.
This is not a storage problem. The same models hold 32 pairs at 0.99 or better, so their memory is eight times bigger than this task needs. It is not slow fading either, because the failure is just as bad at the shortest distance and stays that way across four times the length. Every kind of memory rule failed in exactly the same way. The models never learned to read the pairs in the first place.
Some models read text in both directions, so they see the question before they see the text it refers to. That is a trick known to help recurrent language models, and bidirectional denoisers get it automatically. We could not measure any benefit from it. On a harder version where the distractions reuse the original keys, reading both ways caught on in 3 of 10 runs against 1 of 10 for reading one way, and the difference is not significant.
What did hold up is that two layers are required. With one layer both directions collapse, because the answer sits at the very end of the text and is therefore the first thing the backward pass sees. Whatever it learns there has to be carried forward by a second layer before anything can use it. A shaped training schedule brought both directions to 9 of 10 at exact parity.
We carry this half of the work into DreamingGoose →
Keeping a few attention layers alongside a recurrent model is the usual compromise, and it is normally justified by one of three things. Two of them get weaker here. A model that cannot recall without a convolution can be fixed by adding the convolution, which costs no memory. The retrieval wall that sits well inside the model's capacity is a training problem, fixed by changing which distances the model sees during training.
Three reasons survive. Loads larger than the memory can hold, lengths past the training schedule's limit, and questions whose shape the model cannot anticipate in advance. A decision about whether to keep attention layers, made after adding the filter and fixing the training, is a smaller decision than the one usually being made.
These are small controlled experiments on synthetic tasks, at hidden sizes of 32 and 128. The effect sizes belong to a tightly constrained memory budget, and when memory is plentiful everything solves everything. Our rebuilt Mamba-2 under-trains the official optimised kernel, so margins against real Mamba are softer than the comparisons among our own models, and most of that gap turned out to be learning rate.
What we are claiming is a breakdown of recall into parts, with a number attached to each. The two-layer requirement in the bidirectional result is offered as an explanation rather than demonstrated.