Hardware

How NVIDIA Groq 3 LPX Unlocks Ultrafast Interactivity at Long Context on NVIDIA Vera Rubin

NVIDIA Groq 3 LPX, when paired with the Vera Rubin NVL72 platform, delivers a record‑setting 3,431 output tokens/second on the Artificial Analysis 100K context benchmark using Gemma 4 31B, demonstrati

DGX agentarticle
hardwarenvidia-developer

NVIDIA Groq 3 LPX, when paired with the Vera Rubin NVL72 platform, delivers a record‑setting 3,431 output tokens/second on the Artificial Analysis 100K context benchmark using Gemma 4 31B, demonstrating ultrafast interactivity for long‑context models without sacrificing precision or quality. Its performance is achieved through deterministic compiler‑scheduled workload planning, fine‑grained computation–communication overlap, and preplanned chip‑to‑chip networking that reduces first‑bit latency and enables effective tensor parallelism even at small batch sizes. Benchmarks confirm consistent high throughput across general agentic and coding tasks (e.g., 4,767 median tokens/second on SPEED‑Bench) and support multiple co‑execution configurations—such as prefill‑decode disaggregation and speculative external drafter decoding—to scale to multi‑trillion parameter models.

Related

Source: NVIDIA Developer | 2026-08-24

Loading related sources…