Hardware

Fast LapSum: Exact Differentiable Top-k at Million Scale

arXiv:2608.06912v1 Announce Type: new Abstract: The top-k operation is a fundamental building block of modern sparse computation, enabling token routing, expert activation, memory selection, and atten

DGX agentpaper
hardwarearxiv-cs-ai

arXiv:2608.06912v1 Announce Type: new Abstract: The top-k operation is a fundamental building block of modern sparse computation, enabling token routing, expert activation, memory selection, and attention pruning. Yet standard hard top-k blocks gradients, while existing continuous (soft) relaxations remain too costly for large-scale models. We introduce Fast LapSum, an exact-budget soft top-k primitive whose GPU solver runs in linear time after sorting. Unlike prior linear-time methods such as DFTopK, which relax the normalization constraint, Fast LapSum is, to our knowledge, the first method to preserve an exact selection mass of k while remaining fully differentiable end-to-end. Our solver combines a linear-time threshold computation with an analytical vector--Jacobian product, and for extreme scales employs probabilistic bracketing to sort only the uncertain middle band of kernel-noised scores. The resulting overhead is almost negligible: the solver processes 10^6, 10^7, and 10^8 scores in 0.41, 1.15, and 5.23,ms, respectively. This makes exact soft top-k practical for sparse routing, retrieval, and large-scale optimization. We demonstrate Fast LapSum on two demanding applications operating over millions of coordinates inside the training loop: generating megapixel sparse adversarial examples with an exact soft budget of {sim}0.02% of an image's pixels, achieving an order-of-magnitude speedup over state-of-the-art methods, and training a fully differentiable sparse image coder from scratch.

Related

Source: arXiv cs.AI | 2026-08-10

Loading related sources…