Research
Planning in entropy-regularized Markov decision processes and games
arXiv:2604.19695v1 Announce Type: new Abstract: We propose SmoothCruiser, a new planning algorithm for estimating the value function in entropy-regularized Markov decision processes and two-player gam
arXiv:2604.19695v1 Announce Type: new Abstract: We propose SmoothCruiser, a new planning algorithm for estimating the value function in entropy-regularized Markov decision processes and two-player games, given a generative model of the environment. SmoothCruiser makes use of the smoothness of the Bellman operator promoted by the regularization to achieve problem-independent sample complexity of order O~(1/epsilon^4) for a desired accuracy epsilon, whereas for non-regularized settings there are no known algorithms with guaranteed polynomial sample complexity in the worst case.
Related
- Contraction-Aligned Analysis of Soft Bellman Residual Minimization with Weighted Lp-Norm for Markov Decision Problem
- Blazing the trails before beating the path: Sample-efficient Monte-Carlo planning
- Scale-free adaptive planning for deterministic dynamics & discounted rewards
- Provably Efficient Offline-to-Online Value Adaptation with General Function Approximation
- Simulation-Based Optimisation of Batting Order and Bowling Plans in T20 Cricket
Source: arXiv cs.LG | 2026-04-22