A Tool Bottleneck Framework for Clinically-Informed and Interpretable Medical Image Understanding
DGX agentarXiv:2512.21414v2 Announce Type: replace Abstract: Recent tool-use frameworks powered by vision-language models (VLMs) improve image understanding by grounding model predictions with specialized tool