Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots
DGX agentarXiv:2608.09931v1 Announce Type: new Abstract: Self-improvement for multimodal large language models (MLLMs) is typically driven by reward-based methods that provide only coarse scalar feedback. Dist