Model Releases

Running Qwen 3.6 35B A3B-Q8_0 gguf on a cheap radeon 7600 at 18 token/s * update increased to 21 t/s

I also have 64 gb ddr4 ryzen 5600 Using llama.cpp Ubuntu distro Settings are as follows --n-gpu-layers 999 --n-cpu-moe 36 --no-mmap -ctk q8_0 -ctv q8_0 -fa 1 -c 9000 So rebuilt my llama.cpp build to r

DGX agentreddit
model-releasesr-localllama

I also have 64 gb ddr4 ryzen 5600 Using llama.cpp Ubuntu distro Settings are as follows --n-gpu-layers 999 --n-cpu-moe 36 --no-mmap -ctk q8_0 -ctv q8_0 -fa 1 -c 9000 So rebuilt my llama.cpp build to run rocm 7.14 tokens increased to upper 19 token/per second then overclocked the vram to the maximum LACTL will allow now 21 token/s also. Weird bug if I am watching the tokens being generated by llama it drops to 13 tokens per second but window minimized it goes up to 21 tokens per second weird. submitted by /u/Sweaty_Perception655 [link] [comments]

Related

Source: r/LocalLLaMA | 2026-08-12

Loading related sources…