Location-Aware Pretraining for Medical Difference Visual Question Answering
DGX agentarXiv:2603.04950v2 Announce Type: replace-cross Abstract: Differential medical VQA models compare multiple images to identify clinically meaningful changes and rely on vision encoders to capture fine-