UNIQUE: Universal Top-k Sparse Attention for Training-free Inference and Sparsity-aware Training
arXiv:2605.27740v1 Announce Type: new Abstract: Long-context inference in large language models (LLMs) is bottlenecked by the linear growth of the self-attention key-value (KV) cache. Top-k sparse att