How Well Do Vision-Language Models Understand Sequential Driving Scenes? A Sensitivity Study
DGX agentarXiv:2604.06750v2 Announce Type: replace Abstract: Vision-Language Models (VLMs) are increasingly proposed for autonomous driving tasks, yet their performance on sequential driving scenes remains poo