Model Releases

Been tweaking my Qwen 3.8 setup, up to 45+ steady T/ps at 8bit quant. Realised I'm now top T/ps for this model+ctx across all benchmarked M-series chips. Full args linked below, happy to discuss as this was a pain of trial and error.

https://omlx.ai/benchmarks/performance/2pko3m1k - you can expand the raw args, but I have full annotations of what worked and what didn't. I'm now testing the model on acutal coding and haven't seen a

DGX agentreddit
model-releasesr-localllama

https://omlx.ai/benchmarks/performance/2pko3m1k - you can expand the raw args, but I have full annotations of what worked and what didn't. I'm now testing the model on acutal coding and haven't seen any issues with performance vs default suggested vals for the vanilla model. submitted by /u/Adventurous_Cat_1559 [link] [comments]

Source: r/LocalLLaMA | 2026-08-20

Loading related sources…