NestedKV: Nested Memory Routing for Long-Context KV Cache Compression
DGX agentarXiv:2605.26678v1 Announce Type: new Abstract: Long-context language models are limited by the memory footprint of the key-value (KV) cache. Existing training-free KV compression methods usually rank