Safety
Beyond VLAs: How World Action Models Reshape Robot Manipulation
World Action Models (WAMs) use video-based world modeling instead of vision‑language backbones, giving robots a learned physics engine that supports zero‑shot transfer to new tasks, environments, and
World Action Models (WAMs) use video-based world modeling instead of vision‑language backbones, giving robots a learned physics engine that supports zero‑shot transfer to new tasks, environments, and robot embodiments—overcoming the semantic‑only limitation of Vision‑Language‑Action (VLA) policies. NVIDIA’s open Cosmos 3 model—a Mixture‑of‑Transformers trained on hundreds of millions of images, videos, and action samples—provides a robust foundation for WAM‑based robot policies that can be deployed from high‑throughput workstations to real‑time Jetson devices while delivering measurable performance gains. Policies built on Cosmos 3 require far less task‑specific data yet achieve strong physical generalization across diverse robotic platforms and operating conditions.
Related
- DSWAM: A Dual-System World Action Foundation Model for Fine-Grained Robot Manipulation
- ViVa: A Video-Generative Value Model for Robot Reinforcement Learning
- Action Images: End-to-End Policy Learning via Multiview Video Generation
Source: NVIDIA Developer | 2026-08-04