Hardware
Vera Rubin NVL72 vs GB200 NVL72? Inference TCO & Architecture Analysis
Vera Rubin NVL72 is Nvidia’s second‑generation, rack‑scale Oberon architecture that achieves inference gains through extreme co‑design. Early engineering‑sample data from CoreWeave show DeepSeek R1 de
Vera Rubin NVL72 is Nvidia’s second‑generation, rack‑scale Oberon architecture that achieves inference gains through extreme co‑design. Early engineering‑sample data from CoreWeave show DeepSeek R1 delivering 5.4× higher performance per megawatt and 5× higher performance per dollar than GB200 NVL72, with the gap expected to widen as the architecture matures. Nvidia has released the Rubin (SM_107) software stack with CUDA 13.4, upstreamed kernels for PyTorch, vLLM and OpenAI Triton, and introduced a 3‑bit programmable LUT tensor core; the architecture can reuse Blackwell WGMMA kernels, easing software bring‑up, while the subsequent transition to Feynman (SM_140) will require more extensive kernel rewrites.
Source: SemiAnalysis | 2026-07-23