MaskPro: Linear-Space Probabilistic Learning for Strict (N:M)-Sparsity on LLMs
arXiv:2506.12876v2 Announce Type: replace Abstract: The rapid scaling of large language models~(LLMs) has made inference efficiency a primary bottleneck in the practical deployment. To address this, s