Model Releases

Qwen 3.8 2.4T at 288k tokens/s on Nvidia GB300 NVL72

https://developer.nvidia.com/blog/serve-qwen3-8-2-4t-a95b-a-2-4t-parameter-model-with-configurable-reasoning-on-nvidia-gb300-nvl72/ 4k tokens per second per GPU of which there are 72. 350 tokens per s

DGX agentreddit
model-releasesr-localllama

https://developer.nvidia.com/blog/serve-qwen3-8-2-4t-a95b-a-2-4t-parameter-model-with-configurable-reasoning-on-nvidia-gb300-nvl72/ 4k tokens per second per GPU of which there are 72. 350 tokens per second per user "Without additional model tuning, the model achieves a throughput of over 4K tokens per second per GPU and over 350 tokens per second per user on NVIDIA GB300 NVL72 in FP8 precision on Day 0. Further optimizations, including NVFP4 precision, are expected to deliver enhanced performance gains over time. " submitted by /u/RhubarbSimilar1683 [link] [comments]

Related

Source: r/LocalLLaMA | 2026-08-16

Loading related sources…