Spend Bits Where Queries Look: KV Cache Vector Quantization with Attention-Preserving Transforms
DGX agentarXiv:2608.04074v1 Announce Type: cross Abstract: Long-context LLM decoding reads the key-value (KV) cache at every step. Loading it takes longer than computing attention over it, so throughput is ban