Local Ai
We found a way to run GLM-5.2 with full context in vLLM without pruning. - top 32 experts NVFP4 - rest fp3 - intel autoround calibration - M…
We found a way to run GLM-5.2 with full context in vLLM without pruning. - top 32 experts NVFP4 - rest fp3 - intel autoround calibration - MXP8 attention - bf16 routers and embedded sees On 3 DGX spar
We found a way to run GLM-5.2 with full context in vLLM without pruning. - top 32 experts NVFP4 - rest fp3 - intel autoround calibration - MXP8 attention - bf16 routers and embedded sees On 3 DGX spark / quad RTX pro 6000 Credit to community https://huggingface.co/madeby561/GLM-5.2-MXFP8-NVFP4-NF3-Hybrid
Source: Clem Delangue (X) | 2026-07-06