RAP: KV-Cache Compression via RoPE-Aligned Pruning
arXiv:2602.02599v4 Announce Type: replace Abstract: Long-context inference in large language models (LLMs) is bottlenecked by the memory and compute of the key-value (KV) cache. Structured pruning is