Model Releases
Running Qwen 3.6 35B A3B-Q8_0 gguf on a cheap radeon 7600 at 18 token/s * update increased to 21 t/s
I also have 64 gb ddr4 ryzen 5600 Using llama.cpp Ubuntu distro Settings are as follows --n-gpu-layers 999 --n-cpu-moe 36 --no-mmap -ctk q8_0 -ctv q8_0 -fa 1 -c 9000 So rebuilt my llama.cpp build to r
I also have 64 gb ddr4 ryzen 5600 Using llama.cpp Ubuntu distro Settings are as follows --n-gpu-layers 999 --n-cpu-moe 36 --no-mmap -ctk q8_0 -ctv q8_0 -fa 1 -c 9000 So rebuilt my llama.cpp build to run rocm 7.14 tokens increased to upper 19 token/per second then overclocked the vram to the maximum LACTL will allow now 21 token/s also. Weird bug if I am watching the tokens being generated by llama it drops to 13 tokens per second but window minimized it goes up to 21 tokens per second weird. submitted by /u/Sweaty_Perception655 [link] [comments]
Related
- Running Qwen 3.5 35B A3B-Q8_0 gguf on a cheap radeon 7600 at 18 token/s
- Nifer is insane. 700t/s with Qwen 3.6 35B (no thinking). Purpose build for RTX5090. Full 250k context too.
- Hello from 10KM high! - Thanks to Qwen 3.6 35b a3b!
- Running Qwen 3.6 35B MoE (Q4_K_M) on a Zeus (Xiaomi 12 Pro, 12GB RAM)
Source: r/LocalLLaMA | 2026-08-12