xKV: Cross-Layer KV-Cache Compression via Aligned Singular Vector Extraction
DGX agentarXiv:2503.18893v2 Announce Type: replace Abstract: Long-context Large Language Models (LLMs) enable powerful applications but incur high memory costs due to the key-value states (KV-Cache). Recent st