Graph it first! Enabling Reasoning on Long-form Egocentric Videos through Scene Graphs
DGX agentarXiv:2606.25842v1 Announce Type: new Abstract: Existing multi-modal large language models (MLLMs) face significant challenges in processing long video sequences due to strict input token limitations.