Model Releases
Looking at a PowerColor R9700 for Qwen3.8-27B, Q4_K_XL, llama.cpp/Vulkan.
Hi all Looking at a PowerColor R9700 for Qwen3.8-27B, Q4, llama.cpp/Vulkan. AMD's own blog quotes 51.8 tok/s but doesn't say what context length that's at, or whether MTP=2 was holding up. Separately
Hi all Looking at a PowerColor R9700 for Qwen3.8-27B, Q4, llama.cpp/Vulkan. AMD's own blog quotes 51.8 tok/s but doesn't say what context length that's at, or whether MTP=2 was holding up. Separately I've seen 5090 benchmarks showing Qwen3.8 drops hard as context fills - 75 tok/s at 4K down to around 26 tok/s at 64K, worse degradation than Qwen3.6 apparently. Before I buy: has anyone actually run this combo (R9700, Q4, 64K+ context, real workload not a cold 4K bench) and got real sustained token per sec numbers? Also curious if MTP speculative decoding is stable for anyone yet or still causing OOMs/garbage output like the early CUDA reports. Not after best-case marketing numbers - ideally I'm after "here's what I actually get once the context window's half full." Thanks! submitted by /u/BillyQ [link] [comments]
Related
- AMD llama.cpp: reducing MTP buffer overhead gave me 64K → 149K context for Qwen 27B
- 3 days benchmarking most llama.cpp flags on my weird 40gb vram laptop + tb4 egpu setup. Got +70% generation, +40% prefill, 60k more context, and filed a bug in llama around MTP. What I learned.
- Llama.cpp ROCm 7.2->7.14 upgrade, Radeon 780m iGPU benchmarks: ROCm vs Vulkan
Source: r/LocalLLaMA | 2026-08-23