Model Releases

Nifer is insane. 700t/s with Qwen 3.6 35B (no thinking). Purpose build for RTX5090. Full 250k context too.

I just managed to get it running on windows and this thing is fucking insane. I get around 550-720t/s depending on task at hand. Previously to get to such numbers i would have to do batching and agent

DGX agentreddit
model-releasesr-localllama

I just managed to get it running on windows and this thing is fucking insane. I get around 550-720t/s depending on task at hand. Previously to get to such numbers i would have to do batching and agents in parallel. Here it just does single instance at this insane speed. Couple that with No thinking mode and it fucks so hard that it is not even funny. That's pretty much Cerebras speeds. link to git (linux only but you can build it for windows via something like open code and deepseekv4pro to vibe it.) https://github.com/Neroued/ninfer IT's custom build for RTX5090 and only two models Qwen3.6 27b and 35B. submitted by /u/BringTea_666 [link] [comments]

Source: r/LocalLLaMA | 2026-07-27

Loading related sources…