Research
Discrete Tilt Matching
arXiv:2604.18739v1 Announce Type: new Abstract: Masked diffusion large language models (dLLMs) are a promising alternative to autoregressive generation. While reinforcement learning (RL) methods have
arXiv:2604.18739v1 Announce Type: new Abstract: Masked diffusion large language models (dLLMs) are a promising alternative to autoregressive generation. While reinforcement learning (RL) methods have recently been adapted to dLLM fine-tuning, their objectives typically depend on sequence-level marginal likelihoods, which are intractable for masked diffusion models. To address this, we derive Discrete Tilt Matching (DTM), a likelihood-free method that recasts dLLM fine-tuning as state-level matching of local unmasking posteriors under reward tilting. DTM takes the form of a weighted cross-entropy objective with explicit minimizer, and admits control variates that improve training stability. On a synthetic maze-planning task, we analyze how DTM's annealing schedule and control variates affect training stability and prevent mode collapse. At scale, fine-tuning LLaDA-8B-Instruct with DTM yields strong gains on Sudoku and Countdown while remaining competitive on MATH500 and GSM8K.
Related
- NI Sampling: Accelerating Discrete Diffusion Sampling by Token Order Optimization
- Dataset-Level Metrics Attenuate Non-Determinism: A Fine-Grained Non-Determinism Evaluation in Diffusion Language Models
- Lost in Diffusion: Uncovering Hallucination Patterns and Failure Modes in Diffusion Large Language Models
- Neural Continuous-Time Markov Chain: Discrete Diffusion via Decoupled Jump Timing and Direction
- CRoCoDiL: Continuous and Robust Conditioned Diffusion for Language
- Stability-Weighted Decoding for Diffusion Language Models
Source: arXiv cs.LG | 2026-04-22