Model Releases
Qwen 30b MoE - 30tps - 6GB vram - Done!
So, I have been dreaming of getting 17 tokens per second using my RTX 3050 6GB version on a decent context window for Hermes needed above 60k. The hope is that has was a 22GB of DDR 4, hoping they can
So, I have been dreaming of getting 17 tokens per second using my RTX 3050 6GB version on a decent context window for Hermes needed above 60k. The hope is that has was a 22GB of DDR 4, hoping they can take some of those experts and give me room for context. What did I get 10 or less tokens per second. π Not today!! Today I could run it with 90k context with Hermes I had 20-25 tps. And when I changed harness I got even 30-35tps π₯³π₯³π₯³ NOT benchmarks- but actual session generation with context and actual work being done πππ I will come to edit the post and add details. Just wanted to share the joy with anyone out there with a peasant rig like mine π May be someone who does better can also share the positive vibe. Cheers for now ππΎββοΈ submitted by /u/Bakkario [link] [comments]
Related
- BeeLlama.cpp: advanced DFlash & TurboQuant with support of reasoning and vision. Qwen 3.6 27B Q5 with 200k context on 3090, 2-3x faster than baseline (peak 135 tps!)
- 2 x 5070ti Qwen 27B full config / stats
- Dual 3090 setup: 400 pp t/s to 1600 pp t/s on Qwen 3.6 27B... with slightly lower tps.
- 80 tok/sec and 128K context on 12GB VRAM with Qwen3.6 35B A3B and llama.cpp MTP
Source: r/LocalLLaMA | 2026-08-14