VisionFoundry: Teaching VLMs Visual Perception with Synthetic Images
DGX agentarXiv:2604.09531v1 Announce Type: cross Abstract: Vision-language models (VLMs) still struggle with visual perception tasks such as spatial understanding and viewpoint recognition. One plausible contr