Model Releases
Qwen3.8 2.4T UD-Q1_0 - 178 token generation - 11 min 38s - 0.25 tokens/sec
https://preview.redd.it/nqr9nj5028jh1.png?width=1516&format=png&auto=webp&s=47af40a0737429dc65f2e455b87b2db81725cb14 So I wanted to see... is it possible/viable to run this perhaps once in a while som
https://preview.redd.it/nqr9nj5028jh1.png?width=1516&format=png&auto=webp&s=47af40a0737429dc65f2e455b87b2db81725cb14 So I wanted to see... is it possible/viable to run this perhaps once in a while some hard task.. yea... no. Even with dual 5090's and 3 3090's and 96GB system ram.. still far exceeds my vram + ram by double (not even including context, which was smallish 64k). I am using the "special" UD-Q1_0, which the _0 is the uncommon part, requires a llama.cpp fork, no big deal, it's smaller clearly. 397GB, 115GB smaller than the next IQ1_S version. Regardless, i'd say i have above average vram.. and even using this super small quantized version, 0.25 tokens/sec is far below my limit of usable... this isn't even usable at night to run slow through the night, it's far too slow and would never really get anything done. Of course, I think we all knew this wouldn't really be a local runnable model, but it's still great that it's open weights. Can't wait for Qwen 3.8 27b tomorrow. submitted by /u/klicker0 [link] [comments]
Related
- GLM/Qwen Appreciation Post
- Is this real ? Qwen3.6:27b with 128k context fit in 24Gb VRAM ?
- OrangePi AI Studio Pro - Qwen3.5-122B-A10B
Source: r/LocalLLaMA | 2026-08-13