Mesh-Attention: A New Communication-Efficient Distributed Attention with Improved Data Locality
DGX agentarXiv:2512.20968v2 Announce Type: replace-cross Abstract: Distributed attention is essential for scaling large language models (LLMs) to long contexts, yet existing methods either have limited paralle