Messages, Not Tokens: Grounded Coresets for Faithful VLM Compression
arXiv:2608.02134v1 Announce Type: new Abstract: Modern vision language models (VLMs) turn high-resolution images into long sequences of visual tokens. Every token traverses the language decoder and pe