Compressing Sequences in the Latent Embedding Space: K-Token Merging for Large Language Models
DGX agentarXiv:2604.15153v1 Announce Type: new Abstract: Large Language Models (LLMs) incur significant computational and memory costs when processing long prompts, as full self-attention scales quadratically