Hardware
Run DiffusionGemma on NVIDIA for Developer-Ready, High-Throughput Text Generation
DiffusionGemma is an experimental open model built for exceptionally fast text generation that NVIDIA has optimized to run on GeForce RTX GPUs, RTX PRO, and DGX Spark systems. Rather than generating t
DiffusionGemma is an experimental open model built for exceptionally fast text generation that NVIDIA has optimized to run on GeForce RTX GPUs, RTX PRO, and DGX Spark systems. Rather than generating text one word at a time, DiffusionGemma generates multiple words in parallel to output whole blocks of text, enabling low-latency performance for single-user developer workloads. The model achieves up to 4x-5x faster token output on NVIDIA GPUs, generating over 1,000 tokens per second on a single H100.
Source: NVIDIA Developer | 2026-06-10