Model Releases
Critical attention scaling in long-context transformers
arXiv:2510.05554v2 Announce Type: replace Abstract: As large language models scale to longer contexts, attention layers suffer from a fundamental pathology: attention scores collapse toward uniformity
arXiv:2510.05554v2 Announce Type: replace Abstract: As large language models scale to longer contexts, attention layers suffer from a fundamental pathology: attention scores collapse toward uniformity as context length n increases, causing tokens to cluster excessively, a phenomenon known as rank-collapse. While extit{attention scaling} effectively addresses this deficiency by rescaling attention scores with a polylogarithmic factor eta_n, theoretical justification for this approach remains lacking. We analyze a simplified yet tractable model that magnifies the effect of attention scaling. In this model, attention exhibits a phase transition governed by the scaling factor eta_n: insufficient scaling collapses all tokens to a single direction, while excessive scaling reduces attention to identity, thereby eliminating meaningful interactions between tokens. Our main result identifies the critical scaling eta_n asymp log n and provides a rigorous justification for attention scaling in YaRN and Qwen, clarifying why logarithmic scaling maintains sparse, content-adaptive attention at large context lengths.
Related
- LiteTopK: Exploiting the Curse of Dimensionality for a Fused Indexer-TopK Kernel in Long-Context Sparse Attention
- SinkRouter: Sink-Aware Routing for Efficient Long-Context Decoding in Large Language and Multimodal Models
- Token Sample Complexity of Attention
- Augmenting Attention with Exponentially Decaying Memory Improves Query-Aware KV Sparsity
- MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
Source: arXiv cs.LG | 2026-07-31