Towards Grounded GI Endoscopy VQA via Multi-Task Learning on Small VLMs
DGX agentarXiv:2607.27122v1 Announce Type: new Abstract: Gastrointestinal (GI) endoscopic image analysis has shifted from single-label classification toward visual question answering (VQA), where a model must