Model Releases
Unpopular opinion : Qwen 3.8 27b is not an overthinker
Yes it uses a ton more reasoning tokens than 3.6 did But test in on the same tasks with the other chinese models, glm 5.3, deepseek v4 flash and pro, etc it's really similar, and they are needed The r
Yes it uses a ton more reasoning tokens than 3.6 did But test in on the same tasks with the other chinese models, glm 5.3, deepseek v4 flash and pro, etc it's really similar, and they are needed The reality is, we're just frustrated because our hardware do not allow most of us to have 1M context (I know that it's not supported yet) with 150 tps decode Furthermore, if you don't mind the quality drop, you can just add a reasoning budget, it will still be better than 3.6 submitted by /u/sukazu [link] [comments]
Related
- 1BIT Qwen 3.8 2.4T a95b (unsloth iQ1_S) (MEDIUM Reasoning)
- Observations on Muse-Glimmer reasoning traces being noticeably different from qwen / gemma models and questions for you guys
- BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)
Source: r/LocalLLaMA | 2026-08-17