Hardware
NVIDIA Vera Rubin and Blackwell Set a New Standard for Agentic AI Performance per Watt
NVIDIA’s new GPUs in the Vera Rubin and Blackwell families set a new benchmark for agentic‑AI performance per watt on AgentX, an open‑source inference test that models realistic multi‑step, tool‑using
NVIDIA’s new GPUs in the Vera Rubin and Blackwell families set a new benchmark for agentic‑AI performance per watt on AgentX, an open‑source inference test that models realistic multi‑step, tool‑using workflows with long context and MoE execution. On this workload, Vera Rubin NVL72 delivers up to 30× higher throughput‑per‑megawatt than the previous GB300 NVL72, while GB300 itself achieves up to 80× gains over H200 NVL8 for large Mixture‑of‑Experts models such as Kimi K3 2.8T. These improvements stem from system‑level optimizations—including SGLang, TensorRT‑LLM, vLLM MoE runtimes, DeepGEMM kernels, mixed‑precision formats (MXFP4/8), the NVIDIA Dynamo session‑aware stack, and high‑bandwidth NVLink fabric—allowing coordinated, rack‑scale inference.
Related
- Inside NVIDIA Rubin GPU Architecture: Powering the Era of Agentic AI
- NVIDIA Blackwell Sets STAC-AI Record for LLM Inference in Finance
- How the NVIDIA Vera Rubin Platform is Solving Agentic AI’s Scale-Up Problem
- AgentX - InferenceXv3: Does CUDA Moat Hold up in Agentic Inferencing?
Source: NVIDIA Developer | 2026-08-24