Research
Any-Order GPT as Masked Diffusion Model: Decoupling Formulation and Architecture
arXiv:2506.19935v2 Announce Type: replace-cross Abstract: Efficiently scaling Large Language Models (LLMs) necessitates exploring alternatives to dominant autoregressive (AR) methods, with Masked Diff
arXiv:2506.19935v2 Announce Type: replace-cross Abstract: Efficiently scaling Large Language Models (LLMs) necessitates exploring alternatives to dominant autoregressive (AR) methods, with Masked Diffusion Models (MDMs) emerging as candidates. However, comparing AR (typically decoder-only) and MDM (often encoder-only) paradigms is confounded by differing architectures, obscuring true algorithmic and efficiency trade-offs. This research decouples these factors by evaluating MDMs within a decoder-only framework to: (1) Equitably compare MDM (as Any-Order AR) and standard AR paradigms through discrepancies on orders. (2) Investigate MDM architectural impacts on computational efficiency. We show decoder-only MDMs, despite a larger modeling space, can achieve significant inference speedups (sim25imes) and comparable perplexity with techniques like temperature annealing, offering a path to reduced inference compute. This work provides insights for developing more computationally efficient foundation models by disentangling core modeling choices from architectural influences. Code is available at https://github.com/scxue/AO-GPT-MDM.
Related
- Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion
- A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models
- Not All Denoising Steps Are Equal: Model Scheduling for Faster Masked Diffusion Language Models
- LaViDa-R1: Advancing Reasoning for Unified Multimodal Diffusion Language Models
Source: arXiv cs.CV | 2026-09-02