Mixture-of-Parallelisms: Towards Memory-Efficient Training Stack for Mixture-of-Experts Models
arXiv:2607.01844v1 Announce Type: cross Abstract: This paper showcases a memory-efficient training stack for Mixture-of-Experts (MoE) models. It is a training paradigm that combines and specializes va