Driving Intents Amplify Planning-Oriented Reinforcement Learning
DGX agentarXiv:2605.12625v1 Announce Type: cross Abstract: Continuous-action policies trained on a single demonstrated trajectory per scene suffer from mode collapse: samples cluster around the demonstrated ma