Model Releases
We've gotten some great medium sized models lately (DSV4 Flash 0731, Inkling Small, Laguna S 2.1, Step 3.7 Flash) but does anybody else want to see some new 70-80b contenders?
I can run the mediums, but sometimes I want a faster option that's smarter than Qwen 27B/35B. On my hardware I get like 500 to 800 tok/s prefill and 16 to 22 tok/s gen on ~120B class models, which is
I can run the mediums, but sometimes I want a faster option that's smarter than Qwen 27B/35B. On my hardware I get like 500 to 800 tok/s prefill and 16 to 22 tok/s gen on ~120B class models, which is not the worst but it does get a bit annoying on agentic coding tasks. If we could get some new MoE 70-80B models that are smarter than the Qwen 3.6 family, I would be so happy. Double-ish the prefill/gen would make all the difference. Maybe this is my fault for being cheap and building my GPU rig with some V620's but that price-to-VRAM ratio is hard to beat and I couldn't justify spending more than that so here we are. Or does anyone have some tips? I've been using ROCm + llama.cpp -- I tried using -sm tensor to speed things up, but it's slower. And it gets slower and slower as I try to enable more GPUs with it. So I'm just back to layer split. submitted by /u/TheWolfOfWalmart [link] [comments]
Source: r/LocalLLaMA | 2026-07-31