WaveZip: Wavelet-Driven Space-Time Decoupling for Video Token Condensation
DGX agentarXiv:2607.23265v1 Announce Type: new Abstract: Existing Large Vision-Language Models (LVLMs) struggle with long-form video understanding due to the quadratic computational cost of visual tokens. Whil