Model Releases
Best C++ Local Model? (July 24th 2026 Edition :-P)
I apologize that this question has been asked in various flavors over time, but I couldn't find anything in the posts before that matches the options I have. I have a PC and a mac, both available in m
I apologize that this question has been asked in various flavors over time, but I couldn't find anything in the posts before that matches the options I have. I have a PC and a mac, both available in my network for agentic work. My PC specs are: RTX 5080 16 GB VRAM Intel i7-13700K 32 GB RAM My mac specs are: M4 Max 48 GB RAM I have been using unsloth's qwen3.6-35b-a3b on both machines with Kilo Code as my harness. Both of them have been performing subpar with a small (<10 source files) C++ project. On my PC, I get around 80t/s with a context window of 200k. On my mac, I get around 90t/s with a context window of 150k. Am I already at the pinnacle of local models for C++ agentic work? Is my model optimal? My harness optimal? I feel like I'm spinning my wheels with the results I get from my current agent. Thanks in advance. submitted by /u/McFlurriez [link] [comments]
Related
- Will Ollama come out with a non-cloud version of Deepseek-v4 Flash?
- Need guidance for OLLAMA + Claude setup
- Looking for a local alternative to Claude Code + GSD (Running Qwen 2.5 Coder 14B / Ollama)
- llm-server v2 ai-tuning it self now best performance for llama.cpp/ik_llama.cpp (big steps form v1) auto flag optimization
Source: r/ollama | 2026-07-25