S^4R: Selective Sampling, Subspaces, and Sparse Reconstruction for Compressed Long-Context KV Caching
arXiv:2608.00528v1 Announce Type: new Abstract: The growth of context window lengths in Large Language Models (LLMs) significantly enhances their long-context capabilities but incurs prohibitive memor