Local Ai
Listening to @alexocheema from @exolabs talking about running LLMs locally at @aiDotEngineer @swyx 18 months before we get a SOTA LLM at hom…
Alex Cheema is the co-founder of Exo Labs, a startup founded in March 2024 to 'democratize access to AI' through open source multi-device computing clusters. The core product, *exo*, connects al...
Alex Cheema is the co-founder of Exo Labs, a startup founded in March 2024 to "democratize access to AI" through open source multi-device computing clusters. The core product, exo, connects all your devices into an AI cluster, enabling running models larger than would fit on a single device, with performance that improves as more devices are added. At the AI Engineer conference (featuring hosts such as swyx), Cheema discussed the trajectory of running LLMs locally — a vision already being realized: he connected four Mac Mini M4s with a MacBook Pro M4 Max and ran Alibaba's Qwen2.5Coder-32B using Exo's open-source software, at a total cluster cost of about $5,000 — highly cost-effective compared to a single Nvidia H100 GPU costing between $25,000 and $30,000.
Related
- Cool mlx-vlm segmentation demo by @Prince_Canuma at @aiDotEngineer You can do so much with it
- Using Ollama Gemma4 models via OpenWebUI on my phone and it’s been a good experience
- Tried running LLMs locally to save API costs… ended up waiting 13 minutes for ONE response 🤡
- Is the ASUS ROG Flow Z13 with 128GB of Unified Memory (AMD Strix Halo) a good option to run large LLMs (70B+)?
- MLX creator @awnihannun sharing the story of MLX. Apple management called him right after the launch: “why didn’t you tell us this was going…
Source: local-ai