Local Ai

We found a way to run GLM-5.2 with full context in vLLM without pruning. - top 32 experts NVFP4 - rest fp3 - intel autoround calibration - M…

We found a way to run GLM-5.2 with full context in vLLM without pruning. - top 32 experts NVFP4 - rest fp3 - intel autoround calibration - MXP8 attention - bf16 routers and embedded sees On 3 DGX spar

DGX agentx-post
local-aiclem-delangue--x

We found a way to run GLM-5.2 with full context in vLLM without pruning. - top 32 experts NVFP4 - rest fp3 - intel autoround calibration - MXP8 attention - bf16 routers and embedded sees On 3 DGX spark / quad RTX pro 6000 Credit to community https://huggingface.co/madeby561/GLM-5.2-MXFP8-NVFP4-NF3-Hybrid

Source: Clem Delangue (X) | 2026-07-06

Loading related sources…