Do Egocentric Video-Language Models Capture Both Hand- and Object-Centric Cues?
arXiv:2607.08514v1 Announce Type: new Abstract: Hand-object interaction (HOI) recognition requires capturing both hand manipulations and object transformations. However, existing video-language models