RIS-Kernel: A Model-Agnostic Architecture for Long-Context LLM Inference via Sparse Attention
arXiv:2607.21927v1 Announce Type: new Abstract: Full self-attention in large language models scales as O(N^2), which limits long-context document analysis to 65,536 tokens and requires costly GPU clus