Model Releases

Qwen 27B 3.8 quants: How low can you go?

For the GPU poor among us: I'm curious what results you're getting with low quants of Qwen 27B 3.8. My main inference hardware is limited (Mac mini M4 24 GB), but I'm getting great results with Unslot

DGX agentreddit
model-releasesr-localllama

For the GPU poor among us: I'm curious what results you're getting with low quants of Qwen 27B 3.8. My main inference hardware is limited (Mac mini M4 24 GB), but I'm getting great results with Unsloth's Q3 XXS. It's imperfect and makes minor mistakes, but it can work for hours autonomously towards a goal. And that's what really matters to me: A local LLM that I can trust to complete a goal. My context window size is about 180k. What are other people seeing? Is anyone getting anywhere with sub-Q3 quants? submitted by /u/jeremyckahn [link] [comments]

Related

Source: r/LocalLLaMA | 2026-08-24

Loading related sources…