Model Releases
Tested Nemotron 3.5 Lightning locally on coding, Hermes Agent and agentic work
Ran the model with quants (Q5) and MTP by bartowski with llama.cpp server. It takes ~24GB ram running on M5 Pro with 48GB at about 65t/s. On some tasks it was quite the overthinker. Overall, the quali
Ran the model with quants (Q5) and MTP by bartowski with llama.cpp server. It takes ~24GB ram running on M5 Pro with 48GB at about 65t/s. On some tasks it was quite the overthinker. Overall, the quality of the code output was way below what you can expect for the size (but this is somewhat disclosed by the authors and what this model was optimized for). In Hermes Agent, it did very well in both speed and tool calling capabilities. Watch more: https://www.youtube.com/watch?v=I8Ypa3yK91s submitted by /u/curiousily_ [link] [comments]
Related
- Tested Muse Glimmer locally on coding with OpenCode & agentic work
- Qwen 35B-A3B MoE vs 27B dense in local coding tests: ~4× faster, much smaller quality gap than I expected
- Ran DS V4-Flash-0731 Locally on 3xMI50 32GB @ ~15 t/s TG
- 4x 3090, 96gb vram what Model to drive Hermes?
Source: r/LocalLLaMA | 2026-08-12