CropVLM: Learning to Zoom for Fine-Grained Vision-Language Perception
DGX agentarXiv:2511.19820v2 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) often struggle with tasks that require fine-grained image understanding, such as scene-text recognition or docum