Model Releases
A Human-Factors Guided Cognitive Model of Visuospatial Complexity in Embodied Active Vision
arXiv:2608.23572v1 Announce Type: cross Abstract: We propose a novel framework for the analysis of multimodal data -- encompassing visual, auditory, and spatial stimuli -- foregrounding the role of co
arXiv:2608.23572v1 Announce Type: cross Abstract: We propose a novel framework for the analysis of multimodal data -- encompassing visual, auditory, and spatial stimuli -- foregrounding the role of complexity in embodied perception and interaction in dynamic, naturalistic settings. Grounded in theories of embodied cognition and active vision, we argue that embodied perceptual complexity emerges from an agent's dynamic engagement with the environment and must be analyzed holistically, as a combination of qualitative and quantitative attributes pertaining to, for instance, visuospatial and auditory features. Building on previous work on visual complexity, we expand this into a categorization of diverse complexity attributes -- quantitative, structural, dynamic, auditory, and interactional -- that together characterize multimodal complexity. We demonstrate how this model provides a theoretical framework for characterizing aspects of visuospatial complexity and their interactions, specifically in the context of everyday driving. We also discuss practical applications of the proposed model for creating and evaluating benchmark datasets (e.g., in driving) that centralize cognitive human factors, as well as applications aimed at systematically investigating the effect of visuospatial complexity on human active vision from the viewpoint of visual perception research. The proposed framework lays the foundation for automated methods that interpret complexity in 3D dynamic environments from a human-centered perspective, serving as a semantic template for explainable computational analysis of visuospatial complexity with a categorical focus on cognitive human factors.
Related
- GazeVLM: Active Vision via Internal Attention Control for Multimodal Reasoning
- AVI-Bench: Toward Human-like Audio-Visual Intelligence of Omni-MLLMs
- Is Visual Prompting All You Need? Studying VLM Spatial Reasoning under Progressive Visual Scaffolds
- SkillMemo: Expert-guided Skill Memory Framework for Compositional Embodied Manipulation
Source: arXiv cs.AI | 2026-08-26