Research
Platonic Representation Hypothesis on World Models
arXiv:2608.23720v1 Announce Type: new Abstract: World models have demonstrated significant potential for perceiving and simulating complex environments. Despite their strong performance, the fundament
arXiv:2608.23720v1 Announce Type: new Abstract: World models have demonstrated significant potential for perceiving and simulating complex environments. Despite their strong performance, the fundamental nature of their learned representations remains poorly understood. In this paper, we investigate the Platonic Representation Hypothesis within this domain by proposing the Predictive Consistency Assumption: we posit that the optimization of a shared state transition objective acts as a selective pressure that encourages heterogeneous models to converge toward a shared latent structure. Through systematic experiments with the DINO World Model (DINO-WM), in which we vary visual encoders to create heterogeneous models, we find that capable world models evolve toward geometrically similar internal structures. Moreover, via model stitching, we show that the internal features of one world model can be mapped to another with limited performance degradation, providing evidence of functional compatibility. Our findings suggest that the pursuit of predictive consistency can promote shared, transition-compatible latent structure across world models.
Related
- Platonic Representations in the Human Brain: Unsupervised Recovery of Universal Geometry
- World2VLM: Distilling World Model Imagination into VLMs for Dynamic Spatial Reasoning
- Embody4D: A Generalist 4D World Model for Embodied AI
- Compression and Retrieval: Implicit Memory Retrieval for Video World Models
Source: arXiv cs.CV | 2026-08-26