Safety
Minimal Ingredients for Reward Assignment from Expert Demonstrations
arXiv:2506.06793v2 Announce Type: replace-cross Abstract: Reward assignment from scarce demonstrations is a key challenge in both offline and online imitation learning. A common and intuitive strategy
arXiv:2506.06793v2 Announce Type: replace-cross Abstract: Reward assignment from scarce demonstrations is a key challenge in both offline and online imitation learning. A common and intuitive strategy assigns rewards according to how closely learner trajectories match expert demonstrations. Although this principle underlies many existing methods, the core ingredients that drive performance remain systematically underexplored. We therefore ask: what is the minimal structure that reward assignment must encode to achieve effective downstream RL performance across settings? We approach this question along two design axes: proximity approximation and temporal alignment. Across 32 benchmarks spanning offline and online settings, and with three downstream RL algorithms, our empirical findings suggest: (1) In offline regimes, proximity alone captures the reward structure necessary for effective offline RL, while (2) lightweight temporal correspondence provides consistent gains that are modest offline but essential online or in the presence of multiple demonstrations. We further complement our offline results with a lightweight theory characterizing when simple proximity approximation suffices. Overall, these findings advocate algorithmic minimalism in reward design before introducing complex schemes in both offline and online imitation learning.
Related
- Autonomous Learning From Success and Failure: Goal-Conditioned Supervised Learning with Negative Feedback
- Noise-Guided Transport for Imitation Learning
- Hybrid-AIRL: Enhancing Inverse Reinforcement Learning with Supervised Expert Guidance
- Beyond Binary: Turning Partial Success into Dense Verifiable Rewards for Reinforcement Learning in Code Generation
- GRPO is Secretly a Process Reward Model
Source: arXiv cs.AI | 2026-08-10