Model Releases
I feel like I finally graduated.
I finally made the move from LM Studio to vLLM thanks to this post https://www.reddit.com/r/LocalLLaMA/s/NmS9CgHvqz. I may not know what it all means yet but I’m going to start diving into the docs to
I finally made the move from LM Studio to vLLM thanks to this post https://www.reddit.com/r/LocalLLaMA/s/NmS9CgHvqz. I may not know what it all means yet but I’m going to start diving into the docs to learn as much as I can. I’m running an endpoint on each of my 3090s one for chat and one for subagents. This has made qwen3.8’s reasoning bearable because of the increase to 143tok/s. Thank you to Syv-ai. His repo is here https://github.com/syv-ai/qwen38-27b-rtx3090. vLLM feels like I’m finally using my hardware to its full potential, but the craziest thing is my waterblocked GPUs don’t go above 35°C before they were hitting 70°C on hard workflows. Sorry I didn’t have time to ask qwen to write or edit this post for me. tl;dr vLLm it feels good man submitted by /u/Bpthewise [link] [comments]
Related
- 5060ti Chads, vllm updates and nvfp4
- Updated benchmark: Deepseek V4 Flash on SlopCodeBench (local)
- Optimizing Qwen3.6 / Qwen3.8-27B on 16GB VRAM: Complete Benchmark Results and Setup Guide (~30-50tps at 32k to 72k context)
- DeepSeek V4 Flash 0731 IQ2_M benchmark for Dual 3060 and 96GB RAM ≈ 3.5 tok/s.
Source: r/LocalLLaMA | 2026-08-21