DraftExpert: Expansion-Aware Self-Speculative Decoding for End-Device MoE Inference
DGX agentarXiv:2607.24434v1 Announce Type: cross Abstract: Large Mixture-of-Experts (MoE) language models are attractive for end-device deployment because only a small subset of experts is active per token, bu