Safety
4DGS-WAM: Bridging Past and Future with an Object-Centric World Action Model based on 4D Gaussian Splatting
arXiv:2608.25956v1 Announce Type: new Abstract: Current world action models (WAMs) typically operate on 2D visual data. These models can achieve exceptional visual quality, but they lack explicit spat
arXiv:2608.25956v1 Announce Type: new Abstract: Current world action models (WAMs) typically operate on 2D visual data. These models can achieve exceptional visual quality, but they lack explicit spatial structure for individual objects and repeatedly process redundant background content. Although point clouds can represent the world in 3D space, they can be difficult to align and accumulate across viewpoints. In this paper, we leverage an explicit 4D Gaussian Splatting (4DGS) representation that separately models dynamic objects and the static background of a scene. For dynamic objects, we use a policy model to predict future actor actions and a world model to predict transformations of their observed Gaussian splats. The static background need not be regenerated for future states, as much of it has already been observed in past frames. This forms an object-centric world action model, which we name 4DGS-WAM. It lifts 2D observations into a persistent 4D representation so that previously observed static content can be reused during future prediction. Future-state extrapolation can then focus on modeling the evolution of dynamic objects. Experiments on KITTI-MOT evaluate short-horizon prediction and past reconstruction.
Related
- Scene-Agnostic Object-Centric Representation Learning for 3D Gaussian Splatting
- SG-WAM: Self-Guided World Modeling in Geometry-Aware Policy Space
- Robust-WAM: Bridging Generative Pretraining and Semantic Foresight in World-Action Models
- MVISTA-4D: View-Consistent 4D World Model with Test-Time Action Inference for Robotic Manipulation
Source: arXiv cs.CV | 2026-08-27