Model Releases
TQwen 3.8 flash next ud1s on 6gb vram and 16 gb system ram
After getting tired refreshing and searching sub reddits for qwen 3.8 35 a3b i decided to give qwen 3.8 flash next a try . So I dual booted and built llama cpp I am able to hit 6-7 tps constantly insi
After getting tired refreshing and searching sub reddits for qwen 3.8 35 a3b i decided to give qwen 3.8 flash next a try . So I dual booted and built llama cpp I am able to hit 6-7 tps constantly inside ubuntu just with 16 gb vram and 6 gb gpu . Honestly it is fire for 1 bit quant . Will increase the quant variant till I get somewhat decent speed and acceptable results . What quant will be better to try . I don't wanna download delete and redownload the whole day . https://preview.redd.it/vreqm0btn4mh1.png?width=868&format=png&auto=webp&s=5ccd5d6152bd4c7f52782b79e03af5720c219975 submitted by /u/Agitated_Force_9199 [link] [comments]
Related
- Qwen3.8-Flash-Next (UD-IQ4_XS) on 2x RTX 3060 + 7800X3D, from initial 36 tps prefill to 400 tps and other benchmarks (-sm tensor trap) + VRAM/RAM usage
- Are models with N-Gram tables going to completely change the AI race?
- DeepSeek V4 0731 -> Qwen 3.8 Flash -> GLM 5.3 Flash (and back again!)
Source: r/LocalLLaMA | 2026-08-28