World Models in Words: Auditing Physical State-Transition Commitments in Vision-Language Models
arXiv:2605.29585v1 Announce Type: new Abstract: Vision-language models (VLMs) are increasingly used to answer questions about physical scenes, yet most evaluations reduce performance to a final answer