BoxTuning: Directly Injecting the Object Box for Multimodal Model Fine-Tuning
DGX agentarXiv:2604.11136v1 Announce Type: cross Abstract: Object-level spatial-temporal understanding is essential for video question answering, yet existing multimodal large language models (MLLMs) encode fr