Research
Aggregating Visual Information with Optimal Transport for VideoLM Token Compression
arXiv:2608.20473v1 Announce Type: new Abstract: Video language models process videos as dense visual-token sequences with substantial representational redundancy. Compressing these sequences is theref
arXiv:2608.20473v1 Announce Type: new Abstract: Video language models process videos as dense visual-token sequences with substantial representational redundancy. Compressing these sequences is therefore essential for reducing the visual-token burden on language-model decoding. The central challenge is to preserve visual information dispersed across frames under such compression. To this end, we introduce Aggregating Visual Information with Optimal Transport (AVIOT), which casts video token compression as transporting a dense empirical measure of frame observations onto a compact target measure. The resulting source-to-target coupling induces a distribution over source observations for each target support, directly specifying how the compressed video representation is constructed. We further adapt this construction along task and spatial axes. Question conditioning modulates the transport cost between source frames and target supports, while influencing how many supports are allocated to each temporal segment, thereby directing representation capacity toward question-relevant content. At multiple spatial granularities, AVIOT computes region-specific temporal transport plans and adaptively fuses the representations they yield, allowing different regions within the same compact representation to draw from different moments. Evaluations across varying compression ratios show that AVIOT matches or outperforms the uncompressed baseline on multiple video-understanding benchmarks while retaining strong performance at higher compression ratios.
Related
- Rethinking Video Token Compression with a Global Codebook: Learning Once, Compressing Everywhere
- OTT-Vid: Optimal Transport Temporal Token Compression for Video Large Language Models
- EchoPrune: Interpreting Redundancy as Temporal Echoes for Efficient VideoLLMs
- ST-SimDiff: Balancing Spatiotemporal Similarity and Difference for Efficient Video Understanding with MLLMs
- WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation
Source: arXiv cs.CV | 2026-08-24