Model Releases
So. about the speed of Qwen 3.8 27B Q2_K_XL on 3080 12GB
this screenshot is without MTP usage , fully offloaded onto the GPU. using LM STUDIO i wanted to ask you guys if theres a way to make it even faster , as MTP really didnt help and is infact slower due
this screenshot is without MTP usage , fully offloaded onto the GPU. using LM STUDIO i wanted to ask you guys if theres a way to make it even faster , as MTP really didnt help and is infact slower due to vram overflow and that im on 16GB DDR4 which is disgustingly slow to load models on so i depend on my gpu for every model. this is unsloth's GGUF quant submitted by /u/Infinite_Professor79 [link] [comments]
Related
- QWEN 3.8 27B Q8 Quant - Setup Instructions Help
- Will qwen 3.6 fit in 16gb vram like I can do with 3.5 because of the moe architecture?
Source: r/ollama | 2026-08-20