Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth
DGX agentarXiv:2605.18603v1 Announce Type: new Abstract: Vision-Language Models (VLMs) deployed as situated agents in high-resolution visual environments require active perception -- the ability to dynamically