IceCache: Memory-efficient KV-cache Management for Long-Sequence LLMs
DGX agentarXiv:2604.10539v1 Announce Type: cross Abstract: Key-Value (KV) cache plays a crucial role in accelerating inference in large language models (LLMs) by storing intermediate attention states and avoid