LEAP: Learnable End-to-End Adaptive Pruning of Large Language Models
DGX agentarXiv:2605.17289v1 Announce Type: cross Abstract: Unstructured sparsity is now natively accelerated by recent GPU kernels and dataflow hardware, shifting the bottleneck from inference execution to the