Credit the Right Box: Marginal Contribution Assignment for Structured Visual Perception
DGX agentarXiv:2608.01055v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) are increasingly expected to solve structured perception tasks that require visual recognition, language-to-obj