Local Ai
Yes, we are back👑, with 206 tok/s on a single RTX 5090! Amazing Day-0 work from the SGLang team. Give it a try~@sgl_project
Yes, we are back👑, with 206 tok/s on a single RTX 5090! Amazing Day-0 work from the SGLang team. Give it a try~@sgl_project The king of small models is back! Qwen3.8-27B from @Alibaba_Qwen is open sou
Yes, we are back👑, with 206 tok/s on a single RTX 5090! Amazing Day-0 work from the SGLang team. Give it a try~@sgl_project The king of small models is back! Qwen3.8-27B from @Alibaba_Qwen is open source, and Day-0 support is live in SGLang: - 206.1 tok/s decode on a single RTX 5090, with our NVFP4 plus DSpark - 38.28 tok/s decode on DGX Spark Qwen3.8-27B raises the bar again for what a small model ca…
Related
- 🎁A gift for developers: Qwen3.8-27B, running locally on AMD from day zero. Appreciate the work from the AMD team! @AMD
- ⚡️⚡️Run Qwen3.6-27B locally! @UnslothAI
Source: Qwen (X) | 2026-08-14