SpenseGPT: Practical One-shot Pruning Enabling Sparse and Dense GEMMs for LLM Inference
arXiv:2606.10445v1 Announce Type: cross Abstract: Semi-structured 2:4 sparsity is widely supported by modern accelerators, providing up to a 2x theoretical speedup. However, its strict 50% sparsity co