Multipole Semantic Attention: A Fast Approximation of Softmax Attention for Pretraining
DGX agentarXiv:2509.10406v4 Announce Type: replace Abstract: Pretraining transformers on long sequences (entire code repositories, collections of related documents) is bottlenecked by quadratic attention costs