From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models
arXiv:2605.20177v1 Announce Type: new Abstract: Recent advances in vision-language models (VLMs) emphasize long chain-of-thought reasoning; yet, we find that their performance on visual tasks is prima