Research
Compressing Sequences in the Latent Embedding Space: K-Token Merging for Large Language Models
arXiv:2604.15153v1 Announce Type: new Abstract: Large Language Models (LLMs) incur significant computational and memory costs when processing long prompts, as full self-attention scales quadratically
arXiv:2604.15153v1 Announce Type: new Abstract: Large Language Models (LLMs) incur significant computational and memory costs when processing long prompts, as full self-attention scales quadratically with input length. Token compression aims to address this challenge by reducing the number of tokens representing inputs. However, existing prompt-compression approaches primarily operate in token space and overlook inefficiencies in the latent embedding space. In this paper, we propose K-Token Merging, a latent-space compression framework that merges each contiguous block of K token embeddings into a single embedding via a lightweight encoder. The compressed sequence is processed by a LoRA-adapted LLM, while generation remains in the original vocabulary. Experiments on structural reasoning (Textualized Tree), sentiment classification (Amazon Reviews), and code editing (CommitPackFT) show that K-Token Merging lies on the Pareto frontier of performance vs. compression, achieving up to 75% input length reduction with minimal performance degradation.
Related
- Latent-Condensed Transformer for Efficient Long Context Modeling
- E2LLM: Encoder Elongated Large Language Models for Long-Context Understanding and Reasoning
- Adaptive Layer Selection for Layer-Wise Token Pruning in LLM Inference
- DYCP: Dynamic Context Pruning for Long-Form Dialogue with LLMs
- AsyncTLS: Efficient Generative LLM Inference with Asynchronous Two-level Sparse Attention
Source: arXiv cs.CL | 2026-04-17