PBSD: Privileged Bayesian Self-Distillation for Long-Horizon Credit Assignment
arXiv:2606.09348v1 Announce Type: new Abstract: Long-horizon agentic tasks pose a fundamental credit assignment challenge for outcome-base reinforcement learning: trajectory-level rewards verify final