A long conversation costs memory that grows with every token, and that memory is what limits how much context a model can hold. The cache storing it treats every token as equally important. A small number of them decide where the model looks, and the rest are payload. We are building a cache that spends its bytes accordingly.
When a language model reads, it stores a key and a value for every token it has seen, so it can refer back to them later. That store is the KV cache, and it grows in a straight line with the length of the text. At long contexts it is the largest thing in memory, and it is usually what stops you fitting a bigger context onto a given GPU.
Two things are known about that cache, from separate lines of work. First, attention is very uneven. A few tokens carry most of the routing, and most tokens barely matter. Second, the stored vectors are redundant across their own dimensions, which means they can be compressed. Existing methods use one fact or the other.
ASKV uses both, on the reasoning that they apply to different tokens. The tokens that decide where attention goes need to stay exact. The older, quieter ones are where the redundancy lives, and those are the ones worth compressing hard. Compressing everything the same way would throw away the routing information that makes the rest work.
The kernels exist and they work. Decode runs over the compressed cache directly, with no reconstruction step, and matches the reference arithmetic to within 1.43e-6 on the shapes we have tested. The background maintenance path, the page format and the selection machinery are all built.
The measured results so far are mixed, and we would rather say so than round them up. On a small model matrix, ASKV beat a standard FP8 cache but did not beat an uncompressed one. Memory sweeps do find configurations that use fewer active bytes than FP8 across the context lengths we searched. On quality, higher compression ranks reduce the error and do not eliminate it at the ranks we tested.
None of that settles the question, because every quality and serving number we have runs on models far smaller than the ones this is for. The scale claims need 80GB-class hardware, and until those runs exist we are not claiming the performance benefits this design is supposed to deliver.
ASKV will likely take longer to produce the first token than FP16, FP8 or pruning, unless the compression runs asynchronously, is amortized, or the cache is reused across turns. Long prompts already make prefill expensive, and ASKV adds statistics collection, partitioning and a factoring pass on top. The workloads it suits are ones where a compressed cache can be prepared ahead of a request, reused, or moved between machines.
Quantization is already practical in production, and one published method already keeps the first and most recent tokens at full rank while compressing the middle. The question ASKV has to answer is whether choosing fidelity from attention patterns buys enough over those approaches to justify the complexity.