Learning how to Forget: Fine-tuning for Long-Context Sparse Attention
DGX agentarXiv:2608.19920v1 Announce Type: new Abstract: A lot of prior work addressed key-value (KV) cache selection and compression by sparse attention to enable long-context inference for transformer langua