RP-OPSD: Resolution-Privileged On-Policy Self-Distillation for Multimodal Large Language Models
DGX agentarXiv:2607.24447v1 Announce Type: new Abstract: On-Policy Self-Distillation (OPSD) uses privileged information available only to the teacher to provide dense token-level supervision on trajectories ge