← ALL RESEARCH

The memory was
never the problem.

Linear attention and state space models replace attention with a memory of fixed size that updates as the model reads. That is much cheaper than attention and much worse at looking things up, which is why models built this way usually keep some attention layers alongside the recurrence. We tested the recurrent cells directly to find which parts of them are responsible for the shortfall. Most of it comes from a small piece most of them leave out, and from training that never showed them the situation they were tested on.

THE SHORT VERSION

Most large language models use attention, which lets them look back at any word they have already read. That is why they are good at remembering, and it is also why they get expensive. The cost grows with the length of the text.

The alternative is a recurrent layer. Rather than letting every word look at every other word, it reads left to right and folds what it has seen into a memory of fixed size. A linear attention layer, or a state space model such as Mamba, is built from one of these. It is much cheaper than attention and much worse at looking things up. In practice that is why models built this way usually keep some attention layers as well.

The standard explanation is that the fixed size is the flaw. Once a long text has been compressed into a small state, the detail you wanted is gone. That explanation is mostly wrong. The cells are not running out of room. Two things account for nearly all of the gap. The first is a small filter over the last few words. Mamba-2 includes one, Gated DeltaNet does not, and adding it moves recall from 0.52 to 0.98 at no memory cost. The second is the training data. The cells were never shown the hard version of the task, so they never learned to handle it.

The second finding is the one worth pausing on. Nothing changed except the order and difficulty of the training examples. The architecture, the memory and the parameters were identical. A task the model scored 0.021 on, where guessing scores 0.023, became a task it solved perfectly.

HOW WE TESTED IT

Comparing these models is harder than it sounds. Mamba-2 ships with a short convolution, a filter over the last few inputs. Gated DeltaNet does not have one. RWKV-7 has a decay mechanism the others lack. Run them against each other and you learn which one wins, because three differences move at once and you cannot tell them apart.

So we rebuilt every model ourselves in plain PyTorch, gave them all exactly the same amount of memory (1,024 state elements, about 29k parameters), and then switched one part at a time. Rebuilding them also meant we could run Mamba-2 without its optimised kernel, which keeps the comparison fair.

The test itself is simple. The model reads a list of key and value pairs, then a stretch of filler text, then one of the keys. It has to produce the matching value. It is the smallest task that separates models which can look things up from models that cannot.

WHAT WE FOUND
A small filter does most of the work
The convolution is a window over the last few words. Adding one lifts the gated delta-rule model from 0.518 to 0.983, and the diagonal model from 0.151 to 0.592. It is the largest single effect we measured, it works in both families, and it uses no memory. It does not cost anything elsewhere either. On a separate test of tracking sequences of operations, where these models are normally strong, the version with the filter was better at every length we tried. Comparisons that pit a model without a convolution against a stock Mamba are mostly measuring the missing convolution.
+0.47 / +0.44
The clever memory rule only matters when the filter is missing
Gated DeltaNet updates its memory with a rank-1 delta rule, which is a more expressive update than Mamba's diagonal one. It beats the diagonal version by +0.19 at 16 pairs and +0.32 at 32, but only in models with no convolution. Give both the convolution and the advantage collapses to +0.03, at which point a same-size Mamba-2 ties it at 16 pairs and beats it at 32. The effect is real inside a model. It is not a property of the family, which is what earlier comparisons implied.
+0.32 → +0.03
Forgetting costs nothing you can measure
An early three-seed run suggested that gated decay cost 0.32 in accuracy. At twenty seeds it does not reproduce. What shows up instead is that the two models fail differently. The one without decay locks in on 4 of 20 seeds. The one with decay never goes above 0.80.
p = 0.86
Memory test at 32 pairs, same memory budget, ten seeds unless noted. Guessing the most common answer scores 0.038, and the best strategy that never reads a pair scores 0.068. Every headline result carries a control where the pairings are scrambled after training and the original answer is scored. A model that genuinely looks things up has to fall back to guessing when that happens, and ours does.
THE MODEL THAT WINS

Put the pieces together and the best model is a gated delta rule wearing the same small filter Mamba-2 ships. At the hardest memory setting it reaches 0.983, with a worst seed of 0.917 across ten runs, and the scrambled-pairings control drops it to 0.037, which is guessing.

We also ran it against Mamba-2 using its own optimised kernel on the same GPU. At a shared default learning rate the official kernel scores 0.867 and ours 0.967. Tuned separately, both improve, to 0.955 against 1.000. The gap between them is +0.045. Most of the distance between our rebuilt version and real Mamba turned out to be the learning rate. Almost none of it was the architecture.

TWO FAILURES THAT LOOK ALIKE

Ask these models to hold more pairs and they degrade gently. Everything manages 8 pairs, and the weaker models fade out toward 32. That is a real capacity curve, and it behaves the way you would expect.

Ask the same models to find 4 pairs inside a paragraph of similar-looking distractions and they stop dead. All 45 runs landed between 0.014 and 0.023, against a guessing floor of 0.019. Attention scored 1.000 at every length.

This is not a storage problem. The same models hold 32 pairs at 0.99 or better, so their memory is eight times bigger than this task needs. It is not slow fading either, because the failure is just as bad at the shortest distance and stays that way across four times the length. Every kind of memory rule failed in exactly the same way. The models never learned to read the pairs in the first place.

HOW WE FIXED IT
1
Everyone guesses. Every recurrent model sits at chance, at every length, with one layer and with two.
2
Train on easy distances first. During training, vary how far the question sits from the answer instead of always training at the hardest distance. Same model, same budget, same task: 0.021 → 1.000.
3
It works like a lottery. A run either catches on or it never does, so an average over three runs tells you nothing. Ordinary training caught on in 1 of 10 runs. The varied distances caught on in 7 (p=0.02). Adding extra supervision on top changed nothing (6/10).
4
There is a length limit. At 256 tokens no uniformly-trained model caught on, at either width, at four times the training budget and two learning rates. A shaped schedule caught on in 4 of 5 runs against 0 of 9 for the uniform one.
5
Only advance when it is actually working. At 512 tokens the timed schedule collapsed completely (0/6). Letting the distance grow only while the model is measurably succeeding catches on in 6 of 6 runs (p=0.0011).
Both arms n=6 at 512 tokens. This removes a limit the schedule introduced, not one the model has.
READING IN BOTH DIRECTIONS

Some models read text in both directions, so they see the question before they see the text it refers to. That is a trick known to help recurrent language models, and bidirectional denoisers get it automatically. We could not measure any benefit from it. On a harder version where the distractions reuse the original keys, reading both ways caught on in 3 of 10 runs against 1 of 10 for reading one way, and the difference is not significant.

What did hold up is that two layers are required. With one layer both directions collapse, because the answer sits at the very end of the text and is therefore the first thing the backward pass sees. Whatever it learns there has to be carried forward by a second layer before anything can use it. A shaped training schedule brought both directions to 9 of 10 at exact parity.

We carry this half of the work into DreamingGoose →

WHAT IT CHANGES

Keeping a few attention layers alongside a recurrent model is the usual compromise, and it is normally justified by one of three things. Two of them get weaker here. A model that cannot recall without a convolution can be fixed by adding the convolution, which costs no memory. The retrieval wall that sits well inside the model's capacity is a training problem, fixed by changing which distances the model sees during training.

Three reasons survive. Loads larger than the memory can hold, lengths past the training schedule's limit, and questions whose shape the model cannot anticipate in advance. A decision about whether to keep attention layers, made after adding the filter and fixing the training, is a smaller decision than the one usually being made.

SCOPE

These are small controlled experiments on synthetic tasks, at hidden sizes of 32 and 128. The effect sizes belong to a tightly constrained memory budget, and when memory is plentiful everything solves everything. Our rebuilt Mamba-2 under-trains the official optimised kernel, so margins against real Mamba are softer than the comparisons among our own models, and most of that gap turned out to be learning rate.

What we are claiming is a breakdown of recall into parts, with a number attached to each. The two-layer requirement in the bidirectional result is offered as an explanation rather than demonstrated.

J. Boesch and A. Wee. Anatomy of Associative Recall in Fixed-State Recurrences: A Matched-State Decomposition, an Interference Wall, and a Curriculum That Breaks It. arXiv:2609.16183, September 2026. Preprint of preliminary results, not peer reviewed. Code and result JSONs are in the repository ↗.