Model Releases
Qwen3.8-27B thinking xhigh Vs. thinking off - Apple M5 Max
I ran a few benchmarks on my MacBook Pro M5 Max with oMLX: Running Qwen3.8-27B with thinking off has a big effect on output quality. Running it with thinking on xhigh burns 5.5x the tokens and runs 6x
I ran a few benchmarks on my MacBook Pro M5 Max with oMLX: Running Qwen3.8-27B with thinking off has a big effect on output quality. Running it with thinking on xhigh burns 5.5x the tokens and runs 6x longer. The amount of thinking that Qwen3.8-27B puts into its work is really enormous - this has already been discussed a lot. The big impact of turning thinking completely off is huge - quality wise it basically puts the model behind Qwen3.6-35B-A3B and other MoE models that provide way more Tok/s. Details here: Compare Benchmark Runs | llm-bench.io submitted by /u/DerTomsn [link] [comments]
Related
- Qwen 3.8 27B xhigh vs medium small comparison (+ others for fun)
- AA is the reason for Qwen3.8 27B shipped with xhigh
- SOTA Apple Silicon Inference (August 15, 2026)
- Been tweaking my Qwen 3.8 setup, up to 45+ steady T/ps at 8bit quant. Realised I'm now top T/ps for this model+ctx across all benchmarked M-series chips. Full args linked below, happy to discuss as this was a pain of trial and error.
Source: r/LocalLLaMA | 2026-08-29