Model Releases
2 x 5070ti Qwen 27B full config / stats
Following up on yesterday's post about running everyone's faves on 2 x 16gb cards while maximizing performance and KV. Previous post data used abandoned Cu130 VLLM image. Stats here are done on cu129-
Following up on yesterday's post about running everyone's faves on 2 x 16gb cards while maximizing performance and KV. Previous post data used abandoned Cu130 VLLM image. Stats here are done on cu129-nightly. Which has the KV cache connector fixes and performance improvements. Highlights - 2 concurrent threads run comfortably without generation speed loss. Decode went up to 94-87tps 0-120k context, with prefill 4.6k-2.4k. You get 170k GPU KV and extra 246k with 8GB of RAM. Which makes it very comfortable for local agentic work. A link to full compose file with a lot of additional info on memory usage etc. Hopefully the upcoming small Qwen 3.8 will fit into this setup as well! submitted by /u/val_in_tech [link] [comments]
Related
- If anyone is running qwen 9b or 27b or 35b and getting wrong facts while web search, follow this.
- r/DestroyMyGame destroyed me to the void for using AI. I used Qwen 3.6 27B Q8 with MTP for about 20% of this single HTML file physics shooter game. I remember last year being blown away by GLM 4.5 Air being able to write a somewhat coherent HTML webpage.
- BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)
- How well do multiple GPUs scale for LLM inference? (Trying to understand the basics)
Source: r/LocalLLaMA | 2026-08-06