Show Me Examples: Inferring Visual Concepts from Image Sets
DGX agentarXiv:2607.02402v2 Announce Type: replace Abstract: Vision-language models (VLMs) can follow complex textual instructions, yet they struggle to reason from purely visual context. In particular, curren