Local Ai
Ollama on CPU for domain ChatGPT
I am creating a problem in silo wherein there is a custom flavor of centos which our team develops and has debug strings. I have vcenter server where we make machines for QA. There, got a 50 core vcpu
I am creating a problem in silo wherein there is a custom flavor of centos which our team develops and has debug strings. I have vcenter server where we make machines for QA. There, got a 50 core vcpu and 20gb ram with 100gb of storage for ollama + openWebUI Currently I can scale vcpu as it is on prem and we have some capacity but the only downfall is qwen2.5 is the only thing that works. I got lot of juniors who keep asking same question that are documented somewhere. I can bridge it by simply giving them an internal ChatGPT with domain context in knowledge base. I need help in: - how can I build the most optimized and fast ollama based ChatGPT thing? - it should be pure CPU based thing. I won't get funded for GPU - Average users would be 12 with 20-30 query per day on an average considering they would also lean and train themselves Also as it is confidential domain data, are there any free AI token API keys that I can use for better output ? submitted by /u/i_Am_Robot [link] [comments]
Related
- are there memory limit settings that can be changed?
- My Local Ollama Server Specs make sense?
- What is your current Ollama setup, and which models are you actually using daily?
Source: r/ollama | 2026-07-23