MTraining: Distributed Dynamic Sparse Attention for Efficient Ultra-Long Context Training
arXiv:2510.18830v2 Announce Type: replace Abstract: The adoption of long context windows has become a standard feature in Large Language Models (LLMs), as extended contexts significantly enhance their