Model Releases
How to optimise local AI for lots of RAM but not a lot of VRAM?
Im running a Ryzen 7 5700x, a 3080ti (12GB) with 64GB of RAM. I’m still new to Local AI, and I’ve tried it in the past but none of the previous generations of AI have been good enough for my specific
Im running a Ryzen 7 5700x, a 3080ti (12GB) with 64GB of RAM. I’m still new to Local AI, and I’ve tried it in the past but none of the previous generations of AI have been good enough for my specific niche use case. Yesterday I tried qwen 3.8 27b and it looked really promising. However on my 3080ti it offloaded to RAM slightly and turned the model agonisingly slow (unsure of exact decode or output speed). I didn’t mess with any of the config and was running Ollama. Is there anything I can do to take advantage of my RAM and increase speeds? submitted by /u/Top_Drink8324 [link] [comments]
Related
- Deepseek + Ollama + OpenClaw. Fully local. $0. Here's what you actually lose.
- How we optimized a local Llama 3 agent: From 15s latency and 68% accuracy to 4s and 100% (Full E2E Code & Guide)
- Compared qwen3.6, qwen3-coder, and deepseek-coder on three coding benchmarks. All running locally on Ollama
- How to Run Local LLMs: Ollama Install + DeepSeek + Gemma 4 + Qwen Free G...
Source: r/ollama | 2026-08-18