Industry
Training a diffusion LLM from scratch is challenging, and diffusion inference advantages don't show up at training. So we convert instead. B…
Training a diffusion LLM from scratch is challenging, and diffusion inference advantages don't show up at training. So we convert instead. Building on the TiDAR recipe, we take ZAYA1-8B-base and run d
Training a diffusion LLM from scratch is challenging, and diffusion inference advantages don't show up at training. So we convert instead. Building on the TiDAR recipe, we take ZAYA1-8B-base and run diffusion-conversion mid-training plus diffusion-SFT on our existing stack.
Source: Emad Mostaque (X) | 2026-05-14