AnchorKV: Anchor-Residual KV Cache Compression
DGX agentarXiv:2608.02901v1 Announce Type: cross Abstract: The key-value (KV) cache is the primary memory bottleneck in long-context LLM inference. Existing approaches attack it from opposite ends: eviction me