Research

Scientific terms should have precision. If we use the terms VLM, VLA, WAM in an indiscriminate fashion, as is becoming common in robotics, w…

Scientific terms should have precision. If we use the terms VLM, VLA, WAM in an indiscriminate fashion, as is becoming common in robotics, we are not helping clarity in communication. Let's keep the h

DGX agentx-post
researchyann-lecun--x

Scientific terms should have precision. If we use the terms VLM, VLA, WAM in an indiscriminate fashion, as is becoming common in robotics, we are not helping clarity in communication. Let's keep the historical origins of these terms in mind. VLMs arose as multimodal extensions of LLMs-the training was for tasks like VQA (VIsual Question Answering). These capture the static semantics of the scene behind an image. No dynamics. World Models (e.g. @ylecun , Ha & Schmidhuber 2018) on the other hand are primarily dynamics models, which go back to control theory -1960 (Bellman, Kalman etc.) This makes them natural for robotics planning / policies- I am in a state s, what action a should I perform to get to state s'. In classical control, these models were written down a priori by modeling the physics of the system; today we think of them as learned neural networks trained from temporal data e.g. video, robot trajectories. But the concept is the same. We shouldn't mix this concept with VLMs.

Related

Source: Yann LeCun (X) | 2026-08-23

Loading related sources…