Every Cache Entry Earns Its Place: Global Allocation of Resolution and Coverage for KV Cache Compression
arXiv:2608.07001v1 Announce Type: new Abstract: As large language models (LLMs) process increasingly long contexts, KV cache storage and repeated access have become a major bottleneck. Existing KV cac