WINDQuant: Weight-Informed Neural Decision-Making for Global Mixed-Precision LLM Quantization
arXiv:2605.26660v1 Announce Type: new Abstract: Quantization is an effective approach to reduce the memory footprint and inference cost of large language models (LLMs), yet maintaining performance in