TENP: Trapezoidal Expert Neuron Pruning For Mixture-of-Experts
DGX agentarXiv:2606.09885v1 Announce Type: new Abstract: Mixture-of-Experts large language models (LLMs) scale efficiently through sparse activation, yet their deployment is fundamentally constrained by the la