Model Releases
5700, 48GB RAM and a 3090 24Gb. Best OS and framework/model?
Hi everyone, I have a system which I have been using for gaming, R7 5700X, 48GB DDR4, RTX3090 24GB. But I want to use it for Local AI to reduce my reliance on cloud AI providers (mainly usage limits -
Hi everyone, I have a system which I have been using for gaming, R7 5700X, 48GB DDR4, RTX3090 24GB. But I want to use it for Local AI to reduce my reliance on cloud AI providers (mainly usage limits - accepting some quality loss). I have had it setup with Ubuntu Server and was using the machine as a server to connect via WebUI, but for some reason an update broke the NVIDIA drivers and then broke my install so I’m starting again. The question is what platform do I run as there are so many platform to choose from and many differing opinions, I tried Ollama+OpenWebUI and Unsloth Studio. Ollama ran very slow with the models (Gemma 4 and Qwen3.6), also OpenWebUI occasionally was slow with web/MCP, but it was stable. Unsloth however, was quick and search/tool calls worked perfectly but unstable and the models crashed a few times. I’d just like to know, what is everyone else using for this kind of setup, I can’t get a solid sense of what is the go-to setup for this kind of system is (some say Unsloth, or Ollama, or llama.cpp etc). Also what models are people running well on 24Gb VRAM + 48GB RAM? Its primary job is coding/finding info from the web & PDF’s/generating config files. Thank you submitted by /u/NWSpitfire [link] [comments]
Related
- DKV: Open-source KV-cache compression framework for local LLM inference (CLI + technical report)
- MindControl - llama.cpp fork to guide the reasoning process via injection during sampling
- How I built a free, local AI powerhouse in 10 days (Ollama + Gemma 4 + Claude Cowork 3P + Browserless)
- Ollama Qwen3.6:35b randomly stops outputting tokens
Source: r/LocalLLaMA | 2026-07-25