Local Ai
Running a 31B model locally made me realize how insane LLM infra actually is
A Reddit post from r/ollama in which a user shares their experience running a 31B parameter model locally using Ollama, reflecting on the surprisingly demanding hardware and infrastructure requirement
A Reddit post from r/ollama in which a user shares their experience running a 31B parameter model locally using Ollama, reflecting on the surprisingly demanding hardware and infrastructure requirements involved. The post likely discusses insights into VRAM constraints, memory bandwidth bottlenecks, and quantization trade-offs that become apparent when scaling up to larger models — challenges that cloud LLM providers handle invisibly at massive scale. The discussion serves as a firsthand account of how complex production LLM infrastructure truly is, sparked by the hands-on difficulty of replicating even a fraction of it on consumer hardware.
Related
- llama4 108b
- Tried running LLMs locally to save API costs… ended up waiting 13 minutes for ONE response 🤡
- Cline with Ollama on a RTX4090 (24GRAM) and i9 with 64 GRAM
- My 2026 Ollama Setup Guide: What Actually Works Best for Daily Use on Consumer Hardware
- Recommended Model for a 4060ti 8gb and 16gb ram
- Downloading an AI model just to hit an OOM error is the worst. 📉
Source: r/ollama | 2026-04-15