Detector-Empowered Video Large Language Model for Efficient Spatio-Temporal Grounding
arXiv:2512.06673v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) are rapidly expanding from general video understanding to finer-grained understanding such as spatio-tempor