Model Releases

Experiment with Qwen3.8-Flash-Next on NVIDIA GB300 NVL72 for Agentic Coding

Qwen3.8‑Flash‑Next is a multimodal mixture‑of‑experts model with 125 B parameters that employs Gated DeltaNet (compressing historic context) and Qwen Sparse Attention (exact retrieval over full contex

DGX agentarticle
model-releasesnvidia-developer

Qwen3.8‑Flash‑Next is a multimodal mixture‑of‑experts model with 125 B parameters that employs Gated DeltaNet (compressing historic context) and Qwen Sparse Attention (exact retrieval over full context), enabling a native 262,144‑token window extendable to 1M tokens via YaRN. Alibaba benchmarks show up to 7.6× prefill and 4.9× decoding speedups over full attention for 1 M‑token workloads, with NVIDIA GB300 NVL72 delivering >16 K tokens/second per GPU (>200 tokens/second per user) for high‑throughput agentic coding. The model can be fine‑tuned with NVIDIA NeMo AutoModel and RL recipes, and deployed on NVIDIA‑accelerated platforms using open‑source inference engines such as SGLang, vLLM, and TokenSpeed.

Related

Source: NVIDIA Developer | 2026-08-26

Loading related sources…