Applications
Near-Optimal Sample Complexities of Divergence-based S-rectangular Distributionally Robust Reinforcement Learning
arXiv:2505.12202v3 Announce Type: replace Abstract: Distributionally robust reinforcement learning (DR-RL) has recently gained significant attention as a principled approach that addresses discrepanci
arXiv:2505.12202v3 Announce Type: replace Abstract: Distributionally robust reinforcement learning (DR-RL) has recently gained significant attention as a principled approach that addresses discrepancies between training and testing environments. To balance robustness, conservatism, and computational traceability, the literature has introduced DR-RL models with SA-rectangular and S-rectangular adversaries. While most existing statistical analyses focus on SA-rectangular models, owing to their algorithmic simplicity and the optimality of deterministic policies, S-rectangular models more accurately capture distributional discrepancies in many real-world applications and often yield more effective robust randomized policies. In this paper, we study the empirical value iteration algorithm for divergence-based S-rectangular DR-RL and establish near-optimal sample complexity bounds of widetilde{O}(|S||A|(1-gamma)^{-4}arepsilon^{-2}), where arepsilon is the target accuracy, |S| and |A| denote the cardinalities of the state and action spaces, and gamma is the discount factor. To the best of our knowledge, these are the first sample complexity results for divergence-based S-rectangular models that achieve optimal dependence on |S|, |A|, and arepsilon simultaneously. We further validate this theoretical dependence through numerical experiments on a robust inventory control problem and a theoretical worst-case example, demonstrating the fast learning performance of our proposed algorithm.
Related
- Bayesian Inverse Transition Learning: Learning Dynamics From Near-Optimal Trajectories
- Concave Certificates: Geometric Framework for Distributionally Robust Risk and Complexity Analysis
- Reinforcement Learning Using known Invariances
- Transfer Learning for Loan Recovery Prediction under Distribution Shifts with Heterogeneous Feature Spaces
Source: arXiv cs.LG | 2026-04-29