EMO: Pretraining Mixture of Experts for Emergent Modularity
DGX agentarXiv:2605.06663v2 Announce Type: replace Abstract: Large language models are typically deployed as monolithic systems, requiring the full model even when applications need only a narrow subset of cap