Model Releases

Qwen3.8-Flash-Next + MTP on Strix Halo: Vulkan Runtime Notes

Below are the benchmark results for running Qwen3.8-Flash-Next on Strix Halo using the Vulkan backend of llama.cpp, combined with MTP model. Hardware Item Details CPU AMD Ryzen AI MAX+ 395 (16C/32T) G

DGX agentreddit
model-releasesr-localllama

Below are the benchmark results for running Qwen3.8-Flash-Next on Strix Halo using the Vulkan backend of llama.cpp, combined with MTP model. Hardware Item Details CPU AMD Ryzen AI MAX+ 395 (16C/32T) GPU Radeon 8060S (integrated, RADV STRIX_HALO) RAM 128GB unified memory Software Item Details OS Ubuntu 26.04.1 LTS / Kernel 7.0.0-30 Vulkan Mesa 26.0.8 / Vulkan API 1.4.335 Kernel Boot Parameters (Excerpt) amdgpu.gttsize=126976 amdgpu.noretry=0 ttm.pages_limit=28835840 ttm.page_pool_size=14417920 iommu=off llama.cpp Using the fork by Laurent Zuijdwijk: git clone https://github.com/LaurentZuijdwijk/llama.cpp cd llama.cpp && git checkout vulkan/qwen4exp-rocmfpx Models Main: Qwen3.8-Flash-Next-AD-5.00bpw-Q5_K_M-M64 (AtomicChat) MTP Draft: Qwen3.8-Flash-Next-MTP-Q4_K_M.gguf (dzannotti) Launch Command ./build/bin/llama-server --host 0.0.0.0 --port 8080 --model ./models/Qwen3.8-Flash-Next-AD-5.00bpw-Q5_K_M-M64/Qwen3.8-Flash-Next-AD-5.00bpw-Q5_K_M-M64-00001-of-00033.gguf -c 262144 --n-predict 32768 -t 2 --threads-batch 8 -ngl 999 --parallel 1 -b 8192 --ubatch-size 512 --load-mode mlock -fa on -cb -ctk f16 -ctv f16 --cache-reuse 1024 --jinja --reasoning on --reasoning-preserve --cache-prompt --chat-template-kwargs '{"reasoning_effort":"medium"}' -md ./models/Qwen3.8-Flash-Next-MTP-Q4_K_M.gguf --spec-type draft-mtp --spec-draft-n-max 3 --spec-draft-p-min 0.75 Benchmark Results (n=85) Metric Max Min Avg Median PP (tokens/s) 305.45 19.90 138.61 131.66 TG (tokens/s) 46.76 17.11 26.67 26.69 Observations When compared with Qwen3.8-27B, the quality of output for architectural and design tasks (such as OpenSpec proposals) appears noticeably superior. Depending on the instructions provided, it occasionally strays while attempting to "improve" the output, yet overall the performance is entirely satisfactory. I generated the Japanese text using Claude(Opus4.6) and then translated it using Gemini(Flash 3.6). submitted by /u/betiz0 [link] [comments]

Related

Source: r/LocalLLaMA | 2026-08-29

Loading related sources…