FORGE: Frame Orthogonality in Relevance Geometry for Long-Form Video Understanding
DGX agentarXiv:2607.25266v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have enabled long-form video understanding at a scale that was not previously possible. However, the density of