ButterflyMoE: Compression-Scalable Ternary Experts via Structured Butterfly Orbits
DGX agentarXiv:2601.13563v5 Announce Type: replace-cross Abstract: In current Mixture of Experts (MoE) architectures, linear memory scaling is present, the memory grows as the number of experts increases. N in