← ALL RESEARCH

A model can keep what it knows
and lose what it does.

Pretraining a large model costs a fortune, and converting one into a cheaper design is meant to protect that investment. We tested what actually survives the conversion. The knowledge does. The ability to find something inside a passage does not, and the loss is easy to miss, because the converted model still scores well on everything it was trained on.

THE SHORT VERSION

A Transformer spends most of its compute on attention, which lets every word look at every other word. That is powerful and it gets expensive as the text grows. Two cheaper alternatives exist. One swaps attention for a recurrent model that folds the text into a fixed-size memory. The other keeps attention but changes the job from predicting the next word to filling in blanks, which is how diffusion language models work.

Earlier work does one of those at a time. We did both at once, converting Qwen3 models at 1.7B and 8B into recurrent diffusion models with no attention anywhere, in three stages, so we could tell which stage lost what.

Language modelling survived. The converted models match their teacher's next-word choices 0.746 and 0.779 of the time, which is close to the original. Looking things up did not survive at all. On a memory test where the teacher scores around 0.34, the converted models scored 0.000 at both sizes. Not worse. Zero. They had learned to answer in the right shape without ever reading what they were answering about.

That sounds like a hard limit of the approach. It is not. The models are capable of the task, as the companion paper shows, and the ability comes back with a change to training alone. The catch is that it comes back unpredictably, and the interesting result is how to make it reliable.

BACKGROUND

Dream and DiffuLLaMA adapt pretrained Transformers into diffusion models but keep the Transformer. RADLADS distils attention into recurrence but keeps next-word prediction. Neither crosses both axes. Crossing both is what makes this worth measuring, because it is the version you would actually want if you were trying to run a capable model cheaply.

Doing it in stages is the method. Stage one turns attention into recurrence while the objective stays the same. Stage two adds a backward pass so every position can see both sides, which is what a diffusion model needs. Stage three trains it to fill in blanks, and that is where the retrieval work happens.

HOW WE TESTED IT

The memory test is written as a small passage. The model reads a few key and value pairs, then filler text, then one of the keys, and has to produce the matching value. Guessing scores about 0.023.

Before scoring any converted model we check the test on the original teacher. The teacher has to beat guessing by a wide margin and collapse when the pairs are scrambled, which proves the test measures reading rather than a lucky format. Only then do we run the students. That ordering matters, because a test the teacher also fails would prove nothing about the conversion.

Every number below is a rate across runs rather than an average. In this regime a run either learns to retrieve or it never does, and an average hides that completely.

WHAT SURVIVED AND WHAT DID NOT
Knowledge transferred
After the first stage the converted models agree with their teacher on the next word 0.746 of the time at 1.7B and 0.779 at 8B, with perplexity retention of 0.932 at 1.7B. Predicting the next word is what the teacher was trained to do, and the conversion held onto it.
0.746 / 0.779
Looking things up did not
On the memory test the teacher scores between 0.335 and 0.350. The converted models score nothing at either size, and their score does not move when the pairs are scrambled, which is the signature of answering from habit rather than from the text. Training longer on the diffusion objective does not bring it back.
0.000
The ability was there all along
The same kind of cell, trained from scratch on the companion paper's small benchmark, solves 16-pair recall at 0.996. Nothing about the design prevents it. The conversion simply never taught it, which turns the problem from an architecture question into a training one.
0.996
MAKING IT RELIABLE
1.7B, distance on a timer
Training on passages where the question sits at varying distances from the answer brings the ability back. Advancing that distance on a fixed schedule does not work reliably. A run that has not yet learned the easy cases gets dragged to hard ones anyway, and spends the rest of its budget failing. One run in three caught on, and only partly.
1/3 RUNS
1.7B, distance gated on progress
Advance the distance only while the model is measurably succeeding, and never advance past a failure. The same three runs all caught on, at ceiling, with both controls clean. Runs that stall simply train longer where they are, which is what they needed.
3/3 RUNS
1.7B, on real text
The recipe holds on wikitext at every pair count and distance we tried. The telemetry explains why earlier attempts on real text had failed. Real text is slower to get going, so a schedule on a timer runs past it. Slow to start is not the same as unable to learn.
0.955–0.995
8B, same recipe unchanged
Applied without modification, the recipe lifts recall from nothing to a mean of 0.901, and holds up at twice the number of pairs and beyond the distances it trained on. Two of three runs caught on. The one that did not is a timing problem rather than a failure of the method, and we say why below.
MEAN 0.901
Retrieval on the trained set of symbols, per run, at both sizes. One more ingredient matters: the forward half of the model has to stay trainable. With it frozen, the model learns to produce answers in the right shape and never learns to connect them to the text.
THE PROBLEM THAT SURVIVED EVERYTHING

Every model that passed the memory test, at both sizes and on every successful run, scored 0.000 on a version of the test that uses a set of symbols it never saw in training. What the training installs is a circuit over the particular words it practised with, and it does not extend to words it did not.

The obvious explanation is that the model memorised specific pairs, and that is wrong. We trained a version where the pairs were drawn fresh from a pool of 128 symbols on every batch, so no fixed pair could be memorised. It still caught on, and it still scored 0.000 outside its pool. It also handled distances far beyond anything it trained on. What separates the two cases is which words ever took part in training, not which pairs.

We tried three ways of managing the growing symbol pool and each failed for a different reason, all of them running out of training budget rather than hitting a wall in the method. The fix our own diagnosis suggests, starting with a small pool and widening it, is untested. At 8B the widening run never got past its first tier.

This is the finding with the sharpest practical edge. A converted model can look competent on its training distribution and be guessing just outside it, and the usual way of checking a model, measuring how well it predicts text, will not show you that.
A SECOND CONVERSION, 7B AND HALF ATTENTION

Separately we converted a Qwen2.5-Coder-7B teacher under a different objective, one that fills in blocks of text rather than the whole sequence at once. That objective suits instruction following better, and it let us keep some attention. The student replaces 21 of 28 attention layers with recurrent cells and keeps every fourth one, with 2.91B trainable parameters.

The run finished cleanly, 85,449 training steps over 700M tokens followed by 4,000 steps on chat-formatted data, and reached a held-out loss of 2.819. It still models text noticeably worse than its teacher, and we report it here as a working conversion rather than a finished model. Generation quality and instruction following are for a later paper.

Two things we tried did not work, and both are worth reporting. Aligning the student's internal states to the teacher's did nothing, because the layers we copied over already did that job, and it slowed training by ten to fifteen percent. And a learning rate that looked fine for tens of thousands of steps turned out to be too high, failing deep into training in a way a step-by-step check for loss spikes never noticed.

SCOPE

The conversion results rest on three runs per setting, and one for the real-text run. Rates like 3/3 and 2/3 are not statistically tested here, and we do not treat them as if they were. Several of our negative results, including the 8B miss, are consistent with a fixed training budget cutting off a process that starts slowly, rather than with the method failing. They stay open until that budget is made adaptive.

Memory is measured with a single kind of test, written in a format close to what the models trained on. We use fresh pairs and fresh symbols, but the format is similar, and we are not claiming broad retrieval ability from it.

J. Boesch and A. Wee. DreamingGoose: Staged Distillation from Autoregressive Transformers to Bidirectional Recurrent Diffusion Language Models. Preprint, September 2026, not peer reviewed. Companion to Anatomy of Associative Recall in Fixed-State Recurrences ↗, which isolates the memory mechanism this paper tests at scale. Code and result JSONs are in the repository ↗.