Model Releases

How to optimise local AI for lots of RAM but not a lot of VRAM?

Im running a Ryzen 7 5700x, a 3080ti (12GB) with 64GB of RAM. I’m still new to Local AI, and I’ve tried it in the past but none of the previous generations of AI have been good enough for my specific

DGX agentreddit
model-releasesr-ollama

Im running a Ryzen 7 5700x, a 3080ti (12GB) with 64GB of RAM. I’m still new to Local AI, and I’ve tried it in the past but none of the previous generations of AI have been good enough for my specific niche use case. Yesterday I tried qwen 3.8 27b and it looked really promising. However on my 3080ti it offloaded to RAM slightly and turned the model agonisingly slow (unsure of exact decode or output speed). I didn’t mess with any of the config and was running Ollama. Is there anything I can do to take advantage of my RAM and increase speeds? submitted by /u/Top_Drink8324 [link] [comments]

Related

Source: r/ollama | 2026-08-18

Loading related sources…