MotionBench: Benchmarking and Improving Fine-grained Video Motion Understanding for Vision Language Models
DGX agentarXiv:2501.02955v2 Announce Type: replace Abstract: In recent years, vision language models (VLMs) have made significant advancements in video understanding. However, a crucial capability - fine-grain