A Single Deep Preference-Conditioned Policy for Learning Pareto Coverage Sets
DGX agentarXiv:2605.08946v1 Announce Type: new Abstract: Preference-conditioned multi-objective reinforcement learning aims to learn a single policy that captures trade-offs across preferences, but under nonli