OmniOPD: Logit-Free On-Policy Distillation via Speculative Verification
DGX agentarXiv:2606.01476v1 Announce Type: cross Abstract: On-Policy Distillation (OPD) trains a student model on its own generative trajectories under dense token-level feedback from a stronger teacher, mitig