Local Ai
Unsloth Q1-Q2 Qwen3.8-27B with MTP since the unsloth ones don't ship with for the lowest quants
https://huggingface.co/jojohai/Qwen3.8-27B-MTP-graft Tested on Vulkan, the grafting saves RAM compared to using an external file. What I don't guarantee however is the quality of answers. The model is
https://huggingface.co/jojohai/Qwen3.8-27B-MTP-graft Tested on Vulkan, the grafting saves RAM compared to using an external file. What I don't guarantee however is the quality of answers. The model is very braindead with the Thinking off. However, when asking questions about culture in Brittany the thinking helps the model recover some intelligence so please enable the thinking submitted by /u/Nyghtbynger [link] [comments]
Related
- Inkling-Small-276B-12B, effort 'max' VS Qwen3.6-27B
- Bonsai 27B: 1-bit dense LLM running locally in your browser using custom WebGPU kernels
- Intern S2 Mobius
Source: r/LocalLLaMA | 2026-08-23