QEvict: Recoverable Quantized KV Eviction for Attention-Drift-Robust Long-Context Decoding
DGX agentarXiv:2608.05326v1 Announce Type: cross Abstract: Autoregressive large language model inference is increasingly constrained by the memory footprint of the Key-Value (KV) cache. A dominant line of work