Local Ai
Mac mini M4 48GB
The Mac Mini M4 Pro with 48GB unified memory is a popular choice in the local AI community for running large language models via Ollama, as its Apple Silicon architecture makes all 48GB of RAM dire...
The Mac Mini M4 Pro with 48GB unified memory is a popular choice in the local AI community for running large language models via Ollama, as its Apple Silicon architecture makes all 48GB of RAM directly available for model loading — enabling comfortable inference of 70B parameter models (e.g., Llama 3.1 70B) and 32B models at 15–22 tokens/sec. The M4 Pro's ~273 GB/s memory bandwidth directly drives inference speed, and the machine draws only 30–40W under AI load compared to 350W+ for a comparable GPU-based PC. Reddit discussions in r/ollama highlight this configuration as the practical "sweet spot" for local AI, balancing model capability, power efficiency, and cost ($1,799–$2,000 new).
Related
- Is the ASUS ROG Flow Z13 with 128GB of Unified Memory (AMD Strix Halo) a good option to run large LLMs (70B+)?
- Tried running LLMs locally to save API costs… ended up waiting 13 minutes for ONE response 🤡
- Use the Same Model Across Ollama, LM Studio, Jan, and your Favorite Local AI Apps
- ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference
Source: local-ai