Structured Redundancy Modeling for Efficient Visual Token Pruning in High-Resolution MLLMs
arXiv:2607.23046v1 Announce Type: new Abstract: Recent high-resolution Multimodal Large Language Models (MLLMs) generate thousands of visual tokens per input, leading to a visual token explosion that