Relay Buffer Independent Communication over Pooled HBM for Efficient MoE Inference on Ascend
arXiv:2605.06055v2 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) inference requires large-scale token exchange across devices, making dispatch and combine major bottlenecks in both p