Pretraining a large model costs a fortune, and converting one into a cheaper design is meant to protect that investment. We tested what actually survives the conversion. The knowledge does. The ability to find something inside a passage does not, and the loss is easy to miss, because the converted model still scores well on everything it was trained on.
A Transformer spends most of its compute on attention, which lets every word look at every other word. That is powerful and it gets expensive as the text grows. Two cheaper alternatives exist. One swaps attention for a recurrent model that folds the text into a fixed-size memory. The other keeps attention but changes the job from predicting the next word to filling in blanks, which is how diffusion language models work.
Earlier work does one of those at a time. We did both at once, converting Qwen3 models at 1.7B and 8B into recurrent diffusion models with no attention anywhere, in three stages, so we could tell which stage lost what.
Language modelling survived. The converted models match their teacher's next-word choices 0.746 and 0.779 of the time, which is close to the original. Looking things up did not survive at all. On a memory test where the teacher scores around 0.34, the converted models scored 0.000 at both sizes. Not worse. Zero. They had learned to answer in the right shape without ever reading what they were answering about.
That sounds like a hard limit of the approach. It is not. The models are capable of the task, as the companion paper shows, and the ability comes back with a change to training alone. The catch is that it comes back unpredictably, and the interesting result is how to make it reliable.
Dream and DiffuLLaMA adapt pretrained Transformers into diffusion models but keep the Transformer. RADLADS distils attention into recurrence but keeps next-word prediction. Neither crosses both axes. Crossing both is what makes this worth measuring, because it is the version you would actually want if you were trying to run a capable model cheaply.
Doing it in stages is the method. Stage one turns attention into recurrence while the objective stays the same. Stage two adds a backward pass so every position can see both sides, which is what a diffusion model needs. Stage three trains it to fill in blanks, and that is where the retrieval work happens.
The memory test is written as a small passage. The model reads a few key and value pairs, then filler text, then one of the keys, and has to produce the matching value. Guessing scores about 0.023.
Before scoring any converted model we check the test on the original teacher. The teacher has to beat guessing by a wide margin and collapse when the pairs are scrambled, which proves the test measures reading rather than a lucky format. Only then do we run the students. That ordering matters, because a test the teacher also fails would prove nothing about the conversion.
Every number below is a rate across runs rather than an average. In this regime a run either learns to retrieve or it never does, and an average hides that completely.
Every model that passed the memory test, at both sizes and on every successful run, scored 0.000 on a version of the test that uses a set of symbols it never saw in training. What the training installs is a circuit over the particular words it practised with, and it does not extend to words it did not.
The obvious explanation is that the model memorised specific pairs, and that is wrong. We trained a version where the pairs were drawn fresh from a pool of 128 symbols on every batch, so no fixed pair could be memorised. It still caught on, and it still scored 0.000 outside its pool. It also handled distances far beyond anything it trained on. What separates the two cases is which words ever took part in training, not which pairs.
We tried three ways of managing the growing symbol pool and each failed for a different reason, all of them running out of training budget rather than hitting a wall in the method. The fix our own diagnosis suggests, starting with a small pool and widening it, is untested. At 8B the widening run never got past its first tier.
Separately we converted a Qwen2.5-Coder-7B teacher under a different objective, one that fills in blocks of text rather than the whole sequence at once. That objective suits instruction following better, and it let us keep some attention. The student replaces 21 of 28 attention layers with recurrent cells and keeps every fourth one, with 2.91B trainable parameters.
The run finished cleanly, 85,449 training steps over 700M tokens followed by 4,000 steps on chat-formatted data, and reached a held-out loss of 2.819. It still models text noticeably worse than its teacher, and we report it here as a working conversion rather than a finished model. Generation quality and instruction following are for a later paper.
Two things we tried did not work, and both are worth reporting. Aligning the student's internal states to the teacher's did nothing, because the layers we copied over already did that job, and it slowed training by ten to fifteen percent. And a learning rate that looked fine for tens of thousands of steps turned out to be too high, failing deep into training in a way a step-by-step check for loss spikes never noticed.
The conversion results rest on three runs per setting, and one for the real-text run. Rates like 3/3 and 2/3 are not statistically tested here, and we do not treat them as if they were. Several of our negative results, including the 8B miss, are consistent with a fixed training budget cutting off a process that starts slowly, rather than with the method failing. They stay open until that budget is made adaptive.
Memory is measured with a single kind of test, written in a format close to what the models trained on. We use fresh pairs and fresh symbols, but the format is similar, and we are not claiming broad retrieval ability from it.