Teaching the Way, Not the Answer: Privileged Tutoring Distillation for Multimodal Policy Optimization
DGX agentarXiv:2606.07000v1 Announce Type: new Abstract: Recent post-training methods, particularly Reinforcement Learning with Verifiable Rewards (RLVR), have significantly enhanced the reasoning ability of L