Model Releases

Qwen3.8 2.4T UD-Q1_0 - 178 token generation - 11 min 38s - 0.25 tokens/sec

https://preview.redd.it/nqr9nj5028jh1.png?width=1516&format=png&auto=webp&s=47af40a0737429dc65f2e455b87b2db81725cb14 So I wanted to see... is it possible/viable to run this perhaps once in a while som

DGX agentreddit
model-releasesr-localllama

https://preview.redd.it/nqr9nj5028jh1.png?width=1516&format=png&auto=webp&s=47af40a0737429dc65f2e455b87b2db81725cb14 So I wanted to see... is it possible/viable to run this perhaps once in a while some hard task.. yea... no. Even with dual 5090's and 3 3090's and 96GB system ram.. still far exceeds my vram + ram by double (not even including context, which was smallish 64k). I am using the "special" UD-Q1_0, which the _0 is the uncommon part, requires a llama.cpp fork, no big deal, it's smaller clearly. 397GB, 115GB smaller than the next IQ1_S version. Regardless, i'd say i have above average vram.. and even using this super small quantized version, 0.25 tokens/sec is far below my limit of usable... this isn't even usable at night to run slow through the night, it's far too slow and would never really get anything done. Of course, I think we all knew this wouldn't really be a local runnable model, but it's still great that it's open weights. Can't wait for Qwen 3.8 27b tomorrow. submitted by /u/klicker0 [link] [comments]

Related

Source: r/LocalLLaMA | 2026-08-13

Loading related sources…