Model Releases

TQwen 3.8 flash next ud1s on 6gb vram and 16 gb system ram

After getting tired refreshing and searching sub reddits for qwen 3.8 35 a3b i decided to give qwen 3.8 flash next a try . So I dual booted and built llama cpp I am able to hit 6-7 tps constantly insi

DGX agentreddit
model-releasesr-localllama

After getting tired refreshing and searching sub reddits for qwen 3.8 35 a3b i decided to give qwen 3.8 flash next a try . So I dual booted and built llama cpp I am able to hit 6-7 tps constantly inside ubuntu just with 16 gb vram and 6 gb gpu . Honestly it is fire for 1 bit quant . Will increase the quant variant till I get somewhat decent speed and acceptable results . What quant will be better to try . I don't wanna download delete and redownload the whole day . https://preview.redd.it/vreqm0btn4mh1.png?width=868&format=png&auto=webp&s=5ccd5d6152bd4c7f52782b79e03af5720c219975 submitted by /u/Agitated_Force_9199 [link] [comments]

Related

Source: r/LocalLLaMA | 2026-08-28

Loading related sources…