Model Releases

Experiment with Qwen3.8-Flash-Next 176B Model on NVIDIA GB300 NVL72 for Agentic Coding

The Qwen3.8‑Flash‑Next 176 B multimodal MoE model uses a hybrid Gated DeltaNet / Qwen Sparse Attention (QSA) architecture that compresses long‑context history and aggregates tokens into micro‑blocks f

DGX agentarticle
model-releasesnvidia-developer

The Qwen3.8‑Flash‑Next 176 B multimodal MoE model uses a hybrid Gated DeltaNet / Qwen Sparse Attention (QSA) architecture that compresses long‑context history and aggregates tokens into micro‑blocks for block‑level importance estimation, yielding up to 7.6× faster prefill and 4.9× faster decoding versus full attention and 8.6× more prefill throughput than Qwen3.7‑Plus at a 1 M‑token context.
When run on NVIDIA’s GB300 NVL72 pod—which contains 72 Blackwell Ultra GPUs linked with high‑bandwidth NVLink—the model achieves >16 k tokens/second per GPU and scales from local DGX workstations to rack‑scale deployments, supported end‑to‑end by SGLang, vLLM, TensorRT‑LLM, NeMo AutoModel, and NeMo RL for fine‑tuning and inference.
Its native 262 K token context window is extendable to 1 M tokens via YaRN; Alibaba releases the weights as a preview of the upcoming Qwen4 architecture.

Related

Source: NVIDIA Developer | 2026-08-26

Loading related sources…