Local Ai
Smallest model (& tips) for intelligent computer use via Hermes?
Hello, I have a friend who's using various local LLM's like qwen3.6 27B, 35b-a3b, North Mini Code, and qwen2.5-vl-7b (just for vision). They have a use case where they're trying to have an LLM drive a
Hello, I have a friend who's using various local LLM's like qwen3.6 27B, 35b-a3b, North Mini Code, and qwen2.5-vl-7b (just for vision). They have a use case where they're trying to have an LLM drive an actual machine via hermes' computer_use tool and cua_driver to click through and navigate a complex web app that has many different pages / slides on it, and figure out what it needs to do to advance to the next screen (there can be many many possible page configurations, but essentially there's a 'click this, then click next, or drag this here, and click next, or click this , wait, click next' type thing). The issue they're having is that often the model seems to make mistakes with figuring out how to click on various things, like the 'page navigation' steps at the bottom of the screen (clicking page 2, 3,4 ,5, 'next'), clicking on some circle elements on the screen, etc etc. They've had similar issues and performance whether using a small model for vision, like qwen2.5-vl-7b, which they like because it's local and they don't have to spend on api, because they have a limited budget, but also with using an api model like mimo 2.5, they haven't got much better performance out of that. I'm trying to research, if anybody knows a small model that would function better with driving a real machine with mouse/keyboard input via hermes' computer_use function. Or if they have any tips for the same with existing models we may have tried. submitted by /u/NotARedditUser3 [link] [comments]
Related
- Bought a 5090 to escape API fees. Ended up building a mini datacenter. Sound familiar?
- Do people building local LLM rigs track RTX Ada/workstation card prices, or just consumer cards like the 5090?
- Bonsai 27B: 1-bit dense LLM running locally in your browser using custom WebGPU kernels
- 125 tok/s for Qwen3.6 q4xl on 2x 4060ti is insane perf/dollar
- Qwen3.6 27B on dual RTX 5060 Ti 16GB with vLLM: ~60 tok/s, 204k context working
Source: r/LocalLLaMA | 2026-07-30