Recall Before You Rank: Similarity-Guided Top-K Reuse for Efficient Long-Context Attention
arXiv:2607.27692v1 Announce Type: new Abstract: Top-K sparse attention reduces the cost of Softmax and value aggregation by attending to only a small subset of key--value (KV) entries. However, identi