Q-Fold: Query-Aware Focus-Context Spatio-Temporal Folding for Long Video Understanding
arXiv:2606.12125v1 Announce Type: new Abstract: Long-video understanding remains challenging for multimodal large language models, because temporally extended videos often contain thousands of frames