Peng's Q(lambda) for Conservative Value Estimation in Offline Reinforcement Learning
DGX agentarXiv:2605.14779v1 Announce Type: new Abstract: We propose a model-free offline multi-step reinforcement learning (RL) algorithm, Conservative Peng's Q(lambda) (CPQL). Our algorithm adapts the Peng's