Revise, Don't Freeze: Sampler-Matched Training for Self-Correcting Masked Diffusion Language Models
arXiv:2606.01026v1 Announce Type: new Abstract: Masked diffusion language models (MDLMs) re-predict every position at each denoising step, but standard samplers commit tokens once revealed, leaving th