DV-SFT: Direct Vision Supervision for Fine-Grained Visual Understanding
DGX agentarXiv:2605.26656v1 Announce Type: new Abstract: Multimodal large language models are typically trained end-to-end to predict ground-truth answers, yet supervision signals are applied exclusively to te