DistillCache: KL-Guided Adaptive KV-Cache Eviction for Memory-Efficient LLM Inference
DGX agentarXiv:2608.08878v1 Announce Type: cross Abstract: Transformer-based large language models (LLMs) achieve strong performance across many tasks, but their Key-Value (KV) cache grows linearly with sequen