Restoring the Sweet Spot: Pass-Rate Weighted Self-Distillation for LLM Reasoning
DGX agentarXiv:2605.27765v1 Announce Type: cross Abstract: Self-Distillation Policy Optimization (SDPO) provides dense token-level credit assignment for reinforcement learning with large language models by lev