Applications
Autonomy-of-Heads: Data-Free Sparse Attention from Frozen Query-Key Geometry
arXiv:2608.06849v1 Announce Type: cross Abstract: Long-context LLM inference is bottlenecked by quadratic attention computation and growing KV-cache costs. Existing sparse attention and KV-compression
arXiv:2608.06849v1 Announce Type: cross Abstract: Long-context LLM inference is bottlenecked by quadratic attention computation and growing KV-cache costs. Existing sparse attention and KV-compression methods typically decide which tokens or heads to preserve from runtime attention scores, observation windows, calibration prompts, or learned gates, making head diagnosis input-dependent and costly to deploy. We propose Autonomy-of-Heads (AoH), a data-free method that identifies retrieval and streaming heads from the spectral geometry of query-key projections. AoH defines the kernel attention operator M_h = W_K^{hop}W_Q^h and uses its effective-rank as a weight-space measure of head function: concentrated spectra indicate a small number of dominant query-key matching directions and are associated with retrieval heads, whereas diffuse spectra indicate the absence of a dominant global matching direction and are associated with streaming heads. We further derive an efficient d_ext{head}-dimensional computation that avoids constructing the full d_ext{model}imes d_ext{model} matrix. We conducted extensive experiments across models demonstrating that at 50% sparsity, AoH retains 96.5% of Full Attention performance on average while reducing prefill and decode latency by up to 41.4% and 66.0%, respectively, and KV-cache memory by 50.0% at 256K tokens.
Related
- MATCH: Modulating Attention via In-Context Retrieval for Long-Context Transformers
- KVServe: Service-Aware KV Cache Compression for Communication-Efficient Disaggregated LLM Serving
- Sparser Block-Sparse Attention via Token Permutation
- Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps
Source: arXiv cs.AI | 2026-08-10