Model Releases
At a 1M-token context length, QSA’s attention kernel is up to 7.6× faster in prefill and 4.9× faster in decode. With a 90% prefix-cache hit …
At a 1M-token context length, QSA’s attention kernel is up to 7.6× faster in prefill and 4.9× faster in decode. With a 90% prefix-cache hit rate, Qwen3.8-Flash-Next delivers 8.6× the prefill throughpu
At a 1M-token context length, QSA’s attention kernel is up to 7.6× faster in prefill and 4.9× faster in decode. With a 90% prefix-cache hit rate, Qwen3.8-Flash-Next delivers 8.6× the prefill throughput of Qwen3.7-Plus.
Related
- With only 6B active parameters, Qwen3.8-Flash-Next-Base tops 8 of 14 benchmarks, including MMLU-Pro, SuperGPQA, BBH and GSM8K. And it remain…
- Qwen3.7 Max now available in Go - text only - 1M context - smartest model in the Qwen family to date
- Qwen 3.8-Max is live on Modal, with the full 1M context window and a custom DFlash speculator under the hood. Love seeing our launch partner…
- ✅Implicit caching is now live on Qwen3.7-Max — kicks in automatically, no setup needed. ⚡️Faster + cheaper out of the box. Need higher, more…
- One GPU, 1M context, Day-0 ready. Big props to the vLLM team for the seamless integration!👍 Try Qwen3.8-27B on vLLM: @vllm_project https://…
Source: Qwen (X) | 2026-08-26