Model Releases

2 x 5070ti Qwen 27B full config / stats

Following up on yesterday's post about running everyone's faves on 2 x 16gb cards while maximizing performance and KV. Previous post data used abandoned Cu130 VLLM image. Stats here are done on cu129-

DGX agentreddit
model-releasesr-localllama

Following up on yesterday's post about running everyone's faves on 2 x 16gb cards while maximizing performance and KV. Previous post data used abandoned Cu130 VLLM image. Stats here are done on cu129-nightly. Which has the KV cache connector fixes and performance improvements. Highlights - 2 concurrent threads run comfortably without generation speed loss. Decode went up to 94-87tps 0-120k context, with prefill 4.6k-2.4k. You get 170k GPU KV and extra 246k with 8GB of RAM. Which makes it very comfortable for local agentic work. A link to full compose file with a lot of additional info on memory usage etc. Hopefully the upcoming small Qwen 3.8 will fit into this setup as well! submitted by /u/val_in_tech [link] [comments]

Related

Source: r/LocalLLaMA | 2026-08-06

Loading related sources…