Local Ai
glm 5.1 is doing well
GLM-5.1 is Z.ai's next-generation flagship model for agentic engineering, built on a 754-billion parameter Mixture-of-Experts architecture with 40 billion active parameters per token, a 200,000-tok...
GLM-5.1 is Z.ai's next-generation flagship model for agentic engineering, built on a 754-billion parameter Mixture-of-Experts architecture with 40 billion active parameters per token, a 200,000-token context window, and released under the MIT license. It achieves state-of-the-art performance on SWE-Bench Pro (scoring 58.4) and delivers major improvements over its predecessor in coding, agentic tool use, reasoning, and long-horizon tasks — having been trained entirely on Huawei Ascend 910B chips without any Nvidia hardware. Users in the local-AI/Ollama community have noted it runs via llama.cpp and Ollama (cloud-mode), though full local inference of the unquantized model requires substantial enterprise-grade infrastructure, with quantized GGUF versions via Unsloth reducing the size from ~1.65TB down to approximately 220GB.
Related
- Recommended Model for a 4060ti 8gb and 16gb ram
- How do I know if an AI model could work locally on my computer?
- Mac mini M4 48GB
- Should we be optimizing for limited compute instead of more parameters? Thoughts?
- Is the ASUS ROG Flow Z13 with 128GB of Unified Memory (AMD Strix Halo) a good option to run large LLMs (70B+)?
Source: local-ai