Model Releases
Experiment with Qwen3.8-Flash-Next on NVIDIA GB300 NVL72 for Agentic Coding
Qwen3.8‑Flash‑Next is a multimodal mixture‑of‑experts model with 125 B parameters that employs Gated DeltaNet (compressing historic context) and Qwen Sparse Attention (exact retrieval over full contex
Qwen3.8‑Flash‑Next is a multimodal mixture‑of‑experts model with 125 B parameters that employs Gated DeltaNet (compressing historic context) and Qwen Sparse Attention (exact retrieval over full context), enabling a native 262,144‑token window extendable to 1M tokens via YaRN. Alibaba benchmarks show up to 7.6× prefill and 4.9× decoding speedups over full attention for 1 M‑token workloads, with NVIDIA GB300 NVL72 delivering >16 K tokens/second per GPU (>200 tokens/second per user) for high‑throughput agentic coding. The model can be fine‑tuned with NVIDIA NeMo AutoModel and RL recipes, and deployed on NVIDIA‑accelerated platforms using open‑source inference engines such as SGLang, vLLM, and TokenSpeed.
Related
- Experiment with Qwen3.8-Flash-Next 176B Model on NVIDIA GB300 NVL72 for Agentic Coding
- Serve Qwen3.8-2.4T-A95B, a 2.4T-Parameter Model, with Configurable Reasoning on NVIDIA GB300 NVL72
- ⚡ Meet Qwen3.6-35B-A3B:Now Open-Source!🚀🚀 A sparse MoE model, 35B total params, 3B active. Apache 2.0 license. 🔥 Agentic coding on par wi…
Source: NVIDIA Developer | 2026-08-26