Model Releases
M3 16GB running Ollama (Qwen 9B) is extremely slow (10-12 mins per task). Am I doing something wrong?
Hey everyone, I constantly see high praise for M3 and M4 Macs for local LLM inference, even the base/16GB models. However, my experience has been quite different, and I'm trying to figure out if I hav
Hey everyone, I constantly see high praise for M3 and M4 Macs for local LLM inference, even the base/16GB models. However, my experience has been quite different, and I'm trying to figure out if I have a misconfiguration. I have an M3 Mac with 16GB of RAM. I'm using Ollama to run qwen:9b for some basic "second brain" tasks (specifically using Codex or Claude Code integrated with my Obsidian vault). The issue: It is incredibly slow. A single query to look up my notes is taking around 10 to 12 minutes to complete. I know 16GB has its limits, but this feels excessive. Has anyone successfully run a similar setup with Obsidian on a 16GB Mac? What settings, quantization, or context size limits should I be tweaking in Ollama to get the fast performance everyone else seems to be getting? Any advice is appreciated! submitted by /u/RpHeVil [link] [comments]
Source: r/ollama | 2026-08-09