Elastic MoE: Unlocking the Inference-Time Scalability of Mixture-of-Experts
DGX agentarXiv:2509.21892v2 Announce Type: replace-cross Abstract: Mixture-of-Experts (MoE) models typically fix the number of activated experts k at both training and inference. However, real-world deployment