Sparse Attention as a Range Searching Problem: Towards an Inference-Efficient Index for KV Cache
DGX agentarXiv:2605.06763v1 Announce Type: new Abstract: Sparse attention improves LLM inference efficiency by selecting a subset of key-value entries, but at the cost of potential accuracy degradation. In par