Look-Closer-Then-Diagnose: Confidence-Aware Ultrasound VQA via Active Zooming
arXiv:2605.21652v1 Announce Type: new Abstract: Vision-Language Models (VLMs) have significantly advanced medical visual question answering, yet their performance in ultrasound remains suboptimal. In