Saliency-Aware Regularized Quantization Calibration for Large Language Models
DGX agentarXiv:2605.05693v2 Announce Type: replace Abstract: Post-training quantization (PTQ) is an effective approach for deploying large language models (LLMs) under memory and latency constraints. Most exis