From Segments to Scenes: Temporal Understanding in Autonomous Driving via Vision-Language Model
arXiv:2512.05277v3 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) are increasingly deployed as the perception and reasoning backbone of autonomous agents acting in the wild, with