Self-Routing: Parameter-Free Expert Routing from Hidden States
DGX agentarXiv:2604.00421v2 Announce Type: replace Abstract: Mixture-of-Experts (MoE) layers increase model capacity by activating only a small subset of experts per token, and typically rely on a learned rout