Learning to See What You Need: Gaze Attention for Multimodal Large Language Models
DGX agentarXiv:2605.13080v1 Announce Type: new Abstract: When humans describe a visual scene, they do not process the entire image uniformly; instead, they selectively fixate on regions relevant to their inten