Hardware

Inside NVIDIA Rubin GPU Architecture: Powering the Era of Agentic AI

NVIDIA’s Rubin GPU, the core of the Vera Rubin platform, delivers up to 10× the agentic inference throughput per watt compared with previous generations, using 336 billion transistors, 224 SMs, 896 Te

DGX agentarticle
hardwarenvidia-developer

NVIDIA’s Rubin GPU, the core of the Vera Rubin platform, delivers up to 10× the agentic inference throughput per watt compared with previous generations, using 336 billion transistors, 224 SMs, 896 Tensor Cores (including expanded precision), third‑generation Transformer Engine, and 288 GB HBM4 at 22 TB/s bandwidth. Enhanced features such as a Tensor Memory Accelerator, inline descriptor updates, activation sparsity, adaptive compression, and fine‑grained kernel triggering optimize mixture‑of‑experts scaling, long‑context attention, and latency, maximizing tokens/sec and tokens/watt for sustained agentic workloads. The rack‑scale NVL72 chassis incorporates liquid cooling, DSX MaxLPS power smoothing, cable‑free MGX architecture, and hot‑swappable NVLink trays, supporting multitrillion‑parameter models and up to 40 % more GPUs within the same power envelope, thereby enabling resilient, high‑throughput agentic AI supercomputing.

Source: NVIDIA Developer | 2026-07-21

Loading related sources…