Model Releases
Quad R9700 AI Pro with vLLM-Radiance easily reaching 17,6k PP
https://preview.redd.it/74bmvel9b5nh1.png?width=1602&format=png&auto=webp&s=0d0c1adaa016a486ffd97c4c466e980dc611b139 I've only recently started looking deeper into vLLM after running llama.cpp for a g
https://preview.redd.it/74bmvel9b5nh1.png?width=1602&format=png&auto=webp&s=0d0c1adaa016a486ffd97c4c466e980dc611b139 I've only recently started looking deeper into vLLM after running llama.cpp for a good while. Initially vLLM (official repo) was terribly slow on my four R9700s (tried that one with two as well), however after trying radiance everything changed. Prefill 17636 - TG at that time was 36,6 That prefill spike was two agent profiles working on different tasks simultaneously (one is writing a yt-dlp dl/conversion workflow the other is auditing agents (profiles). Best TG i've hit was 106 Tok/s with a 80% MTP 4 acceptance rate. For reference, I'm running a Gigabyte MZ32-AR0 (Rev 1.0), EPYC 7282 and using Hermes with Qwen 3.8 27b fp8 262k ctx - worth noting that one GPU is actually only running by PCIe 4x8, three full 4x16. On that note i'm also happy to say that vLLM-Radiance does work well with a quad setup in my case - nvtop consistently shows 100% usage of the four cards, officially only dual setups are supported. I hope this doesn't count as a low effort post, i just had to share. //E submitted by /u/im_EDEN [link] [comments]
Related
- Qwen3.8-27B IQ3_XXS wrote a correct multilayer TMM on a 16 GB Quadro — after 100 minutes, 3 compactions, and 108k output tokens
- Deepseek V4 Flash is now ~#2 open weight model to Kimi K3 and >50x cheaper
- Question: Why is prefill unbelievably faster in vLLM than other inference engines?
Source: r/LocalLLaMA | 2026-09-02