← ALL RESEARCH

Most of what a model
remembers doesn't matter.

A long conversation costs memory that grows with every token, and that memory is what limits how much context a model can hold. The cache storing it treats every token as equally important. A small number of them decide where the model looks, and the rest are payload. We are building a cache that spends its bytes accordingly.

THE SHORT VERSION

When a language model reads, it stores a key and a value for every token it has seen, so it can refer back to them later. That store is the KV cache, and it grows in a straight line with the length of the text. At long contexts it is the largest thing in memory, and it is usually what stops you fitting a bigger context onto a given GPU.

Two things are known about that cache, from separate lines of work. First, attention is very uneven. A few tokens carry most of the routing, and most tokens barely matter. Second, the stored vectors are redundant across their own dimensions, which means they can be compressed. Existing methods use one fact or the other.

ASKV uses both, on the reasoning that they apply to different tokens. The tokens that decide where attention goes need to stay exact. The older, quieter ones are where the redundancy lives, and those are the ones worth compressing hard. Compressing everything the same way would throw away the routing information that makes the rest work.

HOW IT WORKS
1
Watch where attention goes. While the model reads, we record which tokens and heads actually get looked at. Tokens that are recent, that act as attention sinks, or that are marked as protected by the application stay at full fidelity.
2
Compress the quiet ones. The rest are stored as low-rank approximations along the hidden dimension, which is where the redundancy is. They keep enough to be useful without costing full price.
3
Never rebuild them. Attention runs directly over the mixed cache with custom kernels. Reconstructing a compressed vector before every attention operation would turn this into a storage format rather than something a model can decode with.
4
Do the compression in the background. Collecting statistics and factoring the matrices takes time, and that time cannot be added to the gap between tokens. A worker compresses a snapshot of older entries while decoding continues, and the cache switches over when the new version is ready.
WHERE THE WORK STANDS

The kernels exist and they work. Decode runs over the compressed cache directly, with no reconstruction step, and matches the reference arithmetic to within 1.43e-6 on the shapes we have tested. The background maintenance path, the page format and the selection machinery are all built.

The measured results so far are mixed, and we would rather say so than round them up. On a small model matrix, ASKV beat a standard FP8 cache but did not beat an uncompressed one. Memory sweeps do find configurations that use fewer active bytes than FP8 across the context lengths we searched. On quality, higher compression ranks reduce the error and do not eliminate it at the ranks we tested.

None of that settles the question, because every quality and serving number we have runs on models far smaller than the ones this is for. The scale claims need 80GB-class hardware, and until those runs exist we are not claiming the performance benefits this design is supposed to deliver.

Every measured result so far comes from models far smaller than the ones this design targets. They establish that the kernels work, not that the approach pays off.
WHAT THE DESIGN HAS TO RESPECT
No reconstruction on the decode path
Specialized kernels run attention directly over the mixed-fidelity cache. Rebuilding cold keys and values before each attention operation would make this a storage codec rather than a cache used during decoding.
Maintenance off the critical path
Factoring and repartitioning must not block token generation. A background worker compresses a snapshot of older sealed pages while decoding continues to append to a full-fidelity or FP8 delta cache.
Copy-on-write, not in-place
At a decode boundary we switch atomically to the worker's new cache manifest. If it is not ready, decoding continues with the current version.
A smaller cache helps nothing if generation has to wait for the compression to finish. Keeping maintenance off the decode path is a requirement of the design, not an implementation detail.
WHAT IT COSTS

ASKV will likely take longer to produce the first token than FP16, FP8 or pruning, unless the compression runs asynchronously, is amortized, or the cache is reused across turns. Long prompts already make prefill expensive, and ASKV adds statistics collection, partitioning and a factoring pass on top. The workloads it suits are ones where a compressed cache can be prepared ahead of a request, reused, or moved between machines.

Quantization is already practical in production, and one published method already keeps the first and most recent tokens at full rank while compressing the middle. The question ASKV has to answer is whether choosing fidelity from attention patterns buys enough over those approaches to justify the complexity.

The design report and reference implementation are complete. Throughput, resident memory and long-context quality at target scale are unmeasured, and are the runs that will decide whether this is worth shipping.