Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps
DGX agentarXiv:2605.16928v1 Announce Type: cross Abstract: Long-context inference in large language models is bottlenecked by the quadratic cost of full attention. Existing efficient alternatives often rely ei