The Yes-Man Syndrome: Benchmarking Abstention in Embodied Robotic Agents
DGX agentarXiv:2605.20544v1 Announce Type: cross Abstract: Vision-language models (VLMs) are used as high-level planners for embodied agents, translating natural language instructions and visual observations i