On the Sample Complexity of Discounted Reinforcement Learning with Optimized Certainty Equivalents
DGX agentarXiv:2605.21763v1 Announce Type: new Abstract: We study risk-sensitive reinforcement learning in finite discounted MDPs, where a generative model of the MDP is assumed to be available. We consider a