VisualFLIP: Do Predictions Depend on Task-Critical Visual Evidence in Multimodal Reasoning?
DGX agentarXiv:2606.07872v1 Announce Type: new Abstract: When a multimodal large language model answers a visual reasoning question correctly, is the prediction actually supported by the task-critical visual e