Model Releases
Best Setup for a 16 GB VRAM + 128 GB RAM System?
Running a 12700k + 5060 Ti 16 gb with 128 gb DDR4 ram and I'm wondering what's the ideal setup to maximize performance. I did some preliminary stuff but I have to admit, I'm still learning and kinda j
Running a 12700k + 5060 Ti 16 gb with 128 gb DDR4 ram and I'm wondering what's the ideal setup to maximize performance. I did some preliminary stuff but I have to admit, I'm still learning and kinda just copy/pasting llama.cpp commands to the terminal haha. It looks like Qwen 3.6 35B A3B seems to be the move, I think at one point I was able to get about 40-60 t/s decode depending on the GGUF quants but not sure which is good to run and what's the best settings. Also tried to the newer Deepseek V4 Flash 0731 and was able to run a Q2 at about 10 t/s which is okay but probably a little too slow for me. Ideally, I want to experiment with a super fast model that I can just talk to and have it make mini edits and have good iteration sessions with while coding. I know we can just SOTAs for agentic stuff (I use pi/opencode with Codex's $20 and it's working fine), so I want to experiment with a more hands ai assisted approach to writing software. I feel like we as a society moved too quickly from ai autocomplete to agents, and I think there's an unexplored gap there that's probably a really nice and solid balance of speed without burnout. Would love y'all's help! submitted by /u/zkelton [link] [comments]
Related
- What’s the community’s favorite benchmark to validate performance?
- 4x 3090, 96gb vram what Model to drive Hermes?
- Serving Deepseek v4 Flash 0731 on 2x DGX Spark — 5-7 GB OS headroom, what would you do to lower VRAM usage and increase OS available RAM?
- Show-off Saturday: Intel Arc B140 build.
Source: r/LocalLLaMA | 2026-08-16