Hardware
How NVIDIA Groq 3 LPX Unlocks Ultrafast Interactivity at Long Context on NVIDIA Vera Rubin
NVIDIA Groq 3 LPX, when paired with the Vera Rubin NVL72 platform, delivers a record‑setting 3,431 output tokens/second on the Artificial Analysis 100K context benchmark using Gemma 4 31B, demonstrati
NVIDIA Groq 3 LPX, when paired with the Vera Rubin NVL72 platform, delivers a record‑setting 3,431 output tokens/second on the Artificial Analysis 100K context benchmark using Gemma 4 31B, demonstrating ultrafast interactivity for long‑context models without sacrificing precision or quality. Its performance is achieved through deterministic compiler‑scheduled workload planning, fine‑grained computation–communication overlap, and preplanned chip‑to‑chip networking that reduces first‑bit latency and enables effective tensor parallelism even at small batch sizes. Benchmarks confirm consistent high throughput across general agentic and coding tasks (e.g., 4,767 median tokens/second on SPEED‑Bench) and support multiple co‑execution configurations—such as prefill‑decode disaggregation and speculative external drafter decoding—to scale to multi‑trillion parameter models.
Related
- Inside NVIDIA Rubin GPU Architecture: Powering the Era of Agentic AI
- NVIDIA Vera Rubin and Blackwell Set a New Standard for Agentic AI Performance per Watt
- How the NVIDIA Vera Rubin Platform is Solving Agentic AI’s Scale-Up Problem
Source: NVIDIA Developer | 2026-08-24