Soft Q(lambda): A multi-step off-policy method for entropy regularised reinforcement learning using eligibility traces
DGX agentarXiv:2604.13780v1 Announce Type: new Abstract: Soft Q-learning has emerged as a versatile model-free method for entropy-regularised reinforcement learning, optimising for returns augmented with a pen