Research
Gaze Behavior in Visual World Experiments Can be Modeled With Off-the-shelf Language-Vision Encoders
arXiv:2608.07282v1 Announce Type: new Abstract: The recent advances in neural language models have also spurred much work in computational psycholinguistics, asking whether neural LMs are also promisi
arXiv:2608.07282v1 Announce Type: new Abstract: The recent advances in neural language models have also spurred much work in computational psycholinguistics, asking whether neural LMs are also promising models of human language processing. However, work has been overwhelmingly focused on the unimodal case of written or spoken language. In contrast, multimodal experimental paradigms, like visual world studies that present participants with both visual and linguistic input simultaneously, have been neglected. In this paper, we present a novel approach that predicts gaze behavior in visual world studies. It does so by combining a simple multi-modal bi-encoder model of the CLIP family with a bimodal attribution method. We demonstrate the ability of this approach to robustly replicate the results of a seminal English visual world study which shows hu- man predictive processing. Remarkably, it does so without a generative architecture and without the need for fine-tuning, despite not being trained for this task.
Related
- From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models
- Vision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language Models
- HAFI-VLM: A Frequency Perspective for Diagnosing and Enhancing Visual Perception in Vision-Language Models
Source: arXiv cs.CL | 2026-08-10