4DVLT: Dynamic Scene Understanding with Worldline-Centered Vision-Language Tracking
arXiv:2606.22631v1 Announce Type: new Abstract: 4D dynamic scene understanding requires grounding language to a persistent worldline that binds identity, metric 3D motion, and synchronized multi-view