OBIT RESEARCH

Exploring more
efficient inference.

We study model conversion, diffusion language modeling, and the systems that make inference more efficient.

Explore the research ↓
PAPERS & PROJECTS

The project pages describe our methods and results, including what we still need to test.

DualGoose: anatomy of associative recall in fixed-state recurrences
Recurrent cells bundle several mechanisms at once, so comparing them end to end can't say which one costs recall. We toggle a short convolution, the transition structure, and decay one at a time at a fixed state budget. The convolution dominates, the rank-1 margin is conditional on its absence, and a retrieval wall that looks architectural turns out to be a training-coverage gap a distance curriculum removes.
PREPRINT →
DreamingGoose: staged distillation from autoregressive Transformers to recurrent diffusion models
Converting a pretrained Transformer across both axes at once keeps what the teacher knows and loses what it can do inside a sequence. The axes are attention becoming recurrence, and next-token prediction becoming iterative denoising. Language modeling transfers at 1.7B and 8B. In-context retrieval comes back at 0.000. A success-gated distance curriculum repairs it, and a token-coverage boundary survives every intervention.
PREPRINT →
ASKV: attention-selected low-rank KV cache compression
A long context costs memory that grows with every token, and the cache treats every token as equally important. Most are not. ASKV keeps the tokens that decide where attention goes at full fidelity and compresses the quiet ones along the hidden dimension, without ever rebuilding them on the decode path. The kernels run. Results at target scale are still unmeasured.
DESIGN REPORT →
RBWO: runtime-bound weight obfuscation for distributed inference
Inference spread across machines we do not own means the workers can read the pieces of the model they run. RBWO stores each piece transformed, so a copied shard will not compute as a model without a secret held somewhere else. It is not private inference and not encryption, and the problem to measure is how usable a stolen model turns out to be.
PROPOSAL →