Understanding Task Transfer in Vision-Language Models
DGX agentarXiv:2511.18787v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) perform well on multimodal benchmarks but lag behind humans and specialized models on visual perception tasks like dep