HART: High-Resolution Annotation-Free Reasoning Technique through a Closed-loop Framework
arXiv:2602.23615v3 Announce Type: replace Abstract: Current Large Multimodal Models (LMMs) struggle with high-resolution visual inputs during the reasoning process, as the number of image tokens incre