Scalable Training of Spatially Grounded 2D Vision-Language Models for Radiology
DGX agentarXiv:2606.20477v2 Announce Type: replace Abstract: We study how to train visually grounded vision-language models (VLMs) for radiology without manual spatial annotations. We introduce RefRad2D, a lar