Model Releases

Serve Qwen3.8-2.4T-A95B, a 2.4T-Parameter Model, with Configurable Reasoning on NVIDIA GB300 NVL72

Alibaba released the open‑weights Qwen3.8‑2.4T‑A95B (Qwen3.8‑Max), a fine‑grained mixture‑of‑experts model with 2.4 trillion parameters, hybrid full‑ and linear‑attention, a one‑million‑token context

DGX agentarticle
model-releasesnvidia-developer

Alibaba released the open‑weights Qwen3.8‑2.4T‑A95B (Qwen3.8‑Max), a fine‑grained mixture‑of‑experts model with 2.4 trillion parameters, hybrid full‑ and linear‑attention, a one‑million‑token context window, and up to 128 K output tokens. NVIDIA demonstrates inference of the model on its GB300 NVL72 GPU, achieving >4k tokens per second per GPU in FP8 precision (≈350 user‑rate tokens/s) with further performance expected from NVFP4 optimizations. The architecture targets demanding agentic workloads such as coding and large‑scale document analysis, requiring data‑center‑scale accelerated compute and co‑design across chips, system architecture, and software for multinode deployments.

Related

Source: NVIDIA Developer | 2026-08-12

Loading related sources…