Industry

Training a diffusion LLM from scratch is challenging, and diffusion inference advantages don't show up at training. So we convert instead. B…

Training a diffusion LLM from scratch is challenging, and diffusion inference advantages don't show up at training. So we convert instead. Building on the TiDAR recipe, we take ZAYA1-8B-base and run d

DGX agentx-post
industryemad-mostaque--x

Training a diffusion LLM from scratch is challenging, and diffusion inference advantages don't show up at training. So we convert instead. Building on the TiDAR recipe, we take ZAYA1-8B-base and run diffusion-conversion mid-training plus diffusion-SFT on our existing stack.

Source: Emad Mostaque (X) | 2026-05-14

Loading related sources…