Meta-Soft: Leveraging Composable Meta-Tokens for Context-Preserving KV Cache Compression
DGX agentarXiv:2605.22337v2 Announce Type: replace Abstract: The KV cache used in large language models has linearly growing time complexity, so LLMs face memory blow-up and reduced decoding efficiency when th