Model Releases
Experiment with Qwen3.8-Flash-Next 176B Model on NVIDIA GB300 NVL72 for Agentic Coding
The Qwen3.8‑Flash‑Next 176 B multimodal MoE model uses a hybrid Gated DeltaNet / Qwen Sparse Attention (QSA) architecture that compresses long‑context history and aggregates tokens into micro‑blocks f
The Qwen3.8‑Flash‑Next 176 B multimodal MoE model uses a hybrid Gated DeltaNet / Qwen Sparse Attention (QSA) architecture that compresses long‑context history and aggregates tokens into micro‑blocks for block‑level importance estimation, yielding up to 7.6× faster prefill and 4.9× faster decoding versus full attention and 8.6× more prefill throughput than Qwen3.7‑Plus at a 1 M‑token context.
When run on NVIDIA’s GB300 NVL72 pod—which contains 72 Blackwell Ultra GPUs linked with high‑bandwidth NVLink—the model achieves >16 k tokens/second per GPU and scales from local DGX workstations to rack‑scale deployments, supported end‑to‑end by SGLang, vLLM, TensorRT‑LLM, NeMo AutoModel, and NeMo RL for fine‑tuning and inference.
Its native 262 K token context window is extendable to 1 M tokens via YaRN; Alibaba releases the weights as a preview of the upcoming Qwen4 architecture.
Related
- Serve Qwen3.8-2.4T-A95B, a 2.4T-Parameter Model, with Configurable Reasoning on NVIDIA GB300 NVL72
- Introducing Qwen3.7-Max from @Alibaba_Qwen, Qwen’s flagship model for the agent era with 1M context and leading performance across agentic c…
- After pushing 1M+ tokens through Qwen 3.8 27B, here is my optimal llama.cpp config for 16GB VRAM (73k Context, Agentic Coding)
- ⚡ Meet Qwen3.6-35B-A3B:Now Open-Source!🚀🚀 A sparse MoE model, 35B total params, 3B active. Apache 2.0 license. 🔥 Agentic coding on par wi…
Source: NVIDIA Developer | 2026-08-26