Model Releases
To run GLM-5.1 locally (744B params, 40B active MoE), full precision needs ~1.65TB disk + enterprise hardware like 8x H200/B200 GPUs. Minimu…
To run GLM-5.1 locally (744B params, 40B active MoE), full precision needs ~1.65TB disk + enterprise hardware like 8x H200/B200 GPUs. Minimum practical setup: Unsloth 2-bit GGUF quant (~220-236GB). Fi
To run GLM-5.1 locally (744B params, 40B active MoE), full precision needs ~1.65TB disk + enterprise hardware like 8x H200/B200 GPUs. Minimum practical setup: Unsloth 2-bit GGUF quant (~220-236GB). Fits on 256GB RAM Mac (unified memory) or PC with 24GB VRAM GPU + 256GB system RAM (MoE offloading via llama.cpp). Download: http://huggingface.co/unsloth/GLM-5.1-GGUF Easier: Use their API (no local hardware). Weights just dropped today—check http://z.ai blog for full guide.
Related
- INCREDIBLE GLM-5.1 weights are now opensource > i’ve had early access to the weights for the past few days > and yeah… this one matters a lo…
- Check out the GLM-5.1 first impressions with Peter on our YouTube https://www.youtube.com/watch?v=f11tVBXWr2g
- A lot of you are asking about the :cloud tag. Here's the deal: GLM-5.1 is a 744B parameter beast. To hit that 54.9 benchmark score without t…
- GLM 5.1 is now LIVE in Atomic Chat SOTA for code & chat – now runs locally with TurboQuant Thanks to @zai_org for open-sourcing this frontie…
- We partnered with @Zai_org to bring GLM-5.1 to Modal. Free to try as an endpoint for the next month. GLM-5.1 further improves upon GLM-5's c…
Source: Zhipu AI (X) | 2026-04-07