Industry
We quantized GLM-5.2 (744B MoE) to 4-bit — and kept its MTP draft head in BF16. → Matches the FP8 release on quality → Runs on 4×H200 instea…
We quantized GLM-5.2 (744B MoE) to 4-bit — and kept its MTP draft head in BF16. → Matches the FP8 release on quality → Runs on 4×H200 instead of 8 → Fastest 4-bit GLM-5.2 at int conc: +69–79% vs AWQ /
We quantized GLM-5.2 (744B MoE) to 4-bit — and kept its MTP draft head in BF16. → Matches the FP8 release on quality → Runs on 4×H200 instead of 8 → Fastest 4-bit GLM-5.2 at int conc: +69–79% vs AWQ / NVFP4 at batch-1, from MTP speculative decoding 👇 http://huggingface.co/canada-quant/GLM-5.2-W4A16-MTP
Source: Clem Delangue (X) | 2026-06-29