UBP2: Uncertainty-Balanced Preference Planning for Efficient Preference-based Reinforcement Learning
DGX agentarXiv:2606.19328v2 Announce Type: replace Abstract: Preference-based RL provides an approach to learning reward models from pairwise comparisons of behaviors, bypassing the need for explicit reward de