Safety
Optimal Reward Shaping: Autonomous Car Parking Case Study
arXiv:2607.23617v1 Announce Type: new Abstract: Designing effective reward functions for model-free reinforcement learning under non-holonomic constraints remains a persistent challenge, often resulti
arXiv:2607.23617v1 Announce Type: new Abstract: Designing effective reward functions for model-free reinforcement learning under non-holonomic constraints remains a persistent challenge, often resulting in severe local minima such as policy paralysis or over-conservative hazard avoidance. In this work, we present a parameterized reward shaping framework featuring coverage-gated alignment feedback, drive-direction switch regularization, and an aligned episode termination mechanism evaluated on an autonomous parallel parking task. Crucially, we show that environmental reward parameters and algorithmic hyperparameters are deeply co-dependent, requiring joint meta-optimization to achieve stable convergence. By employing surrogate-based Bayesian optimization, our co-optimized Deep Q-Network (DQN) agent resolves characteristic control failure modes, significantly outperforming uncalibrated baselines across both success rate and trajectory smoothness.
Related
- Hybrid Energy-Aware Reward Shaping: A Unified Lightweight Physics-Guided Methodology for Policy Optimization
- Continuous-time reinforcement learning: ellipticity enables model-free value function approximation
- Zero-Shot, Safe and Time-Efficient UAV Navigation via Potential-Based Reward Shaping, Control Lyapunov and Barrier Functions
Source: arXiv cs.LG | 2026-07-28