QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding
arXiv:2607.04559v1 Announce Type: new Abstract: The performance of vision-language models (VLMs) in video understanding declines with increasing video duration, as video moments unrelated to the query