From My View to Yours: Learning Egocentric Cues from Exocentric Video using Privileged Egocentric Supervision
arXiv:2501.05711v4 Announce Type: replace Abstract: Vision Language Models (VLMs) have achieved strong performance across a wide range of video understanding tasks. However, their viewpoint-invariant