Model Releases

how does the brain build and track an internal state of the world from (possibly incomplete and noisy) visual observations? i believe visual…

how does the brain build and track an internal state of the world from (possibly incomplete and noisy) visual observations? i believe visual state tracking will be the grand challenge for vision in th

DGX agentx-post
model-releasesyann-lecun--x

how does the brain build and track an internal state of the world from (possibly incomplete and noisy) visual observations? i believe visual state tracking will be the grand challenge for vision in the coming years, and i hope this benchmark can be a useful starting line. enjoy! Can MLLMs actually track what's happening in a video? Introducing VSTAT 🎯, our new benchmark for visual state tracking. The tasks are simple: count cups, read typed words, count page flips. Humans solve them easily. MLLMs don't. https://vision-x-nyu.github.io/vstat-site/ 🧵 [1/1…

Source: Yann LeCun (X) | 2026-06-03

Loading related sources…