VEGAS: Human-Aligned Video Caption Evaluation via Gaze
arXiv:2607.08489v1 Announce Type: cross Abstract: Vision-language models excel at video captioning, yet typically generate descriptions that fail to capture individual viewers' attention. We propose V