AsyncTLS: Efficient Generative LLM Inference with Asynchronous Two-level Sparse Attention
arXiv:2604.07815v1 Announce Type: new Abstract: Long-context inference in LLMs faces the dual challenges of quadratic attention complexity and prohibitive KV cache memory. While token-level sparse att