Industry

We quantized GLM-5.2 (744B MoE) to 4-bit — and kept its MTP draft head in BF16. → Matches the FP8 release on quality → Runs on 4×H200 instea…

We quantized GLM-5.2 (744B MoE) to 4-bit — and kept its MTP draft head in BF16. → Matches the FP8 release on quality → Runs on 4×H200 instead of 8 → Fastest 4-bit GLM-5.2 at int conc: +69–79% vs AWQ /

DGX agentx-post
industryclem-delangue--x

We quantized GLM-5.2 (744B MoE) to 4-bit — and kept its MTP draft head in BF16. → Matches the FP8 release on quality → Runs on 4×H200 instead of 8 → Fastest 4-bit GLM-5.2 at int conc: +69–79% vs AWQ / NVFP4 at batch-1, from MTP speculative decoding 👇 http://huggingface.co/canada-quant/GLM-5.2-W4A16-MTP

Source: Clem Delangue (X) | 2026-06-29

Loading related sources…