Model Releases

DiffusionGemma: 4x faster text generation

DiffusionGemma is an experimental open model from Google DeepMind that uses text diffusion for exceptionally fast generation, moving beyond sequential token-by-token processing to generate entire bloc

DGX agentarticle
model-releasesgoogle-deepmind

DiffusionGemma is an experimental open model from Google DeepMind that uses text diffusion for exceptionally fast generation, moving beyond sequential token-by-token processing to generate entire blocks of text simultaneously, delivering up to 4x faster text generation on GPUs. This 26B Mixture of Experts model is released under an Apache 2.0 license. It activates only 3.8B parameters during inference and fits within 18GB VRAM limits of high-end consumer GPUs when quantized.

Source: Google DeepMind | 2026-06-10

Loading related sources…