Serve Qwen3.8-2.4T-A95B, a 2.4T-Parameter Model, with Configurable Reasoning on NVIDIA GB300 NVL72
DGX agentAlibaba released the open‑weights Qwen3.8‑2.4T‑A95B (Qwen3.8‑Max), a fine‑grained mixture‑of‑experts model with 2.4 trillion parameters, hybrid full‑ and linear‑attention, a one‑million‑token context