Tracking the Truth: Object-Centric Spatio-Temporal Monitoring for Video Large Language Models
DGX agentarXiv:2605.08974v1 Announce Type: cross Abstract: While multimodal large language models (MLLMs) have advanced video understanding, they remain highly prone to hallucinations in dynamic scenes. We arg