Model Releases

Qwen 30b MoE - 30tps - 6GB vram - Done!

So, I have been dreaming of getting 17 tokens per second using my RTX 3050 6GB version on a decent context window for Hermes needed above 60k. The hope is that has was a 22GB of DDR 4, hoping they can

DGX agentreddit
model-releasesr-localllama

So, I have been dreaming of getting 17 tokens per second using my RTX 3050 6GB version on a decent context window for Hermes needed above 60k. The hope is that has was a 22GB of DDR 4, hoping they can take some of those experts and give me room for context. What did I get 10 or less tokens per second. πŸ˜„ Not today!! Today I could run it with 90k context with Hermes I had 20-25 tps. And when I changed harness I got even 30-35tps πŸ₯³πŸ₯³πŸ₯³ NOT benchmarks- but actual session generation with context and actual work being done πŸ˜„πŸ˜„πŸ˜„ I will come to edit the post and add details. Just wanted to share the joy with anyone out there with a peasant rig like mine πŸ˜… May be someone who does better can also share the positive vibe. Cheers for now πŸ™‹πŸΎβ€β™‚οΈ submitted by /u/Bakkario [link] [comments]

Related

Source: r/LocalLLaMA | 2026-08-14

Loading related sources…