Model Releases

we made Qwen 3.8 27b MLX vision quants and compared them against other popular community publishers (lm-studio, lukaskremla, mlx-community and etc)

we made vision mlx quants of qwen3.8 27b (9 builds from 8bit at 29.5 GB down to 3.23bpw DWQ at 11.8 GB) and compared them against other community vision mlx quants from hf (we only compared vision bui

DGX agentreddit
model-releasesr-localllama

we made vision mlx quants of qwen3.8 27b (9 builds from 8bit at 29.5 GB down to 3.23bpw DWQ at 11.8 GB) and compared them against other community vision mlx quants from hf (we only compared vision builds) the layout comes out of a clipping search we wrote for mlx and on top of that we wanted to try distillation, so the low-bit files are dwq builds, trained against the bf16 model as the teacher. there are still a couple of percent lying around in almost every one of them, so we will keep playing with these quants and posting the results we downloaded every vision mlx build of qwen3.8 we could find and measured all of them ourselves in the same harness: 1x H200 NVL mlx-vlm at ctx 4096 our own held-out text, never used for calibration every file scored against the same reference top-1 agreement (noise floor 0.084%) our 11.8 GB file is at 70.32% top-1 while the other two files under 12 GB are at 53% and 43% the 11.8 GB one runs on a 16 GB macbook if you raise the wired limit with sudo sysctl iogpu.wired_limit_mb=13000. you only get 13-14 GB out of the 16, so the context will be small, but it is enough to play with 😉 everything else is for 24 GB and up explore our collection on hf https://huggingface.co/collections/AtomicChat/qwen-38-27b - there you will find more detailed description of our quantization method and extra benchmarks also you can run every mlx build in our local ai open source app https://atomic.chat - i'm cofounder, so feel free to ask any questions and share your feedback! submitted by /u/Top-Eye-8104 [link] [comments]

Related

Source: r/LocalLLaMA | 2026-08-26

Loading related sources…