Research
Towards Faster Language Model Inference Using Mixture-of-Experts Flow Matching
arXiv:2604.15009v1 Announce Type: cross Abstract: Flow matching retains the generation quality of diffusion models while enabling substantially faster inference, making it a compelling paradigm for ge
arXiv:2604.15009v1 Announce Type: cross Abstract: Flow matching retains the generation quality of diffusion models while enabling substantially faster inference, making it a compelling paradigm for generative modeling. However, when applied to language modeling, it exhibits fundamental limitations in representing complex latent distributions with irregular geometries, such as anisotropy and multimodality. To address these challenges, we propose a mixture-of-experts flow matching (MoE-FM) framework, which captures complex global transport geometries in latent space by decomposing them into locally specialized vector fields. Building on MoE-FM, we develop a non-autoregressive (NAR) language modeling approach, named YAN, instantiated with both Transformer and Mamba architectures. Across multiple downstream tasks, YAN achieves generation quality on par with both autoregressive (AR) and diffusion-based NAR language models, while requiring as few as three sampling steps. This yields a 40imes speedup over AR baselines and up to a 10^3imes speedup over diffusion language models, demonstrating substantial efficiency advantages for language modeling.
Related
- Adjoint Matching through the Lens of the Stochastic Maximum Principle in Optimal Control
- Not All Denoising Steps Are Equal: Model Scheduling for Faster Masked Diffusion Language Models
- EvoESAP: Non-Uniform Expert Pruning for Sparse MoE
- MoE Routing Testbed: Studying Expert Specialization and Routing Behavior at Small Scale
- BezierFlow: Learning Bezier Stochastic Interpolant Schedulers for Few-Step Generation
Source: arXiv cs.LG | 2026-04-17