Agents
Human Cognition in Machines: A Unified Perspective of World Models
arXiv:2604.16592v1 Announce Type: cross Abstract: This comprehensive report distinguishes prior works by the cognitive functions they innovate. Many works claim an almost 'human-like' cognitive capabi
arXiv:2604.16592v1 Announce Type: cross Abstract: This comprehensive report distinguishes prior works by the cognitive functions they innovate. Many works claim an almost "human-like" cognitive capability in their world models. To evaluate these claims requires a proper grounding in first principles in Cognitive Architecture Theory (CAT). We present a conceptual unified framework for world models that fully incorporates all the cognitive functions associated with CAT (i.e. memory, perception, language, reasoning, imagining, motivation, and meta-cognition) and identify gaps in the research as a guide for future states of the art. In particular, we find that motivation (especially intrinsic motivation) and meta-cognition remain drastically under-researched, and we propose concrete directions informed by active inference and global workspace theory to address them. We further introduce Epistemic World Models, a new category encompassing agent frameworks for scientific discovery that operate over structured knowledge. Our taxonomy, applied across video, embodied, and epistemic world models, suggests research directions where prior taxonomies have not.
Related
- MultiWorld: Scalable Multi-Agent Multi-View Video World Models
- Act Wisely: Cultivating Meta-Cognitive Tool Use in Agentic Multimodal Models
- CogDriver: Integrating Cognitive Inertia for Temporally Coherent Planning in Autonomous Driving
- XEmbodied: A Foundation Model with Enhanced Geometric and Physical Cues for Large-Scale Embodied Environments
- MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding
Source: arXiv cs.CV | 2026-04-21