Safety

Beyond Pessimism: Offline Learning in KL-regularized Games

arXiv:2604.06738v1 Announce Type: cross Abstract: We study offline learning in KL-regularized two-player zero-sum games, where policies are optimized under a KL constraint to a fixed reference policy.

DGX agentpaper
safetyarxiv-cs-lg

arXiv:2604.06738v1 Announce Type: cross Abstract: We study offline learning in KL-regularized two-player zero-sum games, where policies are optimized under a KL constraint to a fixed reference policy. Prior work relies on pessimistic value estimation to handle distribution shift, yielding only widetilde{O}(1/sqrt n) statistical rates. We develop a new pessimism-free algorithm and analytical framework for KL-regularized games, built on the smoothness of KL-regularized best responses and a stability property of the Nash equilibrium induced by skew symmetry. This yields the first widetilde{O}(1/n) sample complexity bound for offline learning in KL-regularized zero-sum games, achieved entirely without pessimism. We further propose an efficient self-play policy optimization algorithm and prove that, with a number of iterations linear in the sample size, it achieves the same fast widetilde{O}(1/n) statistical rate as the minimax estimator.

Related

Source: arXiv cs.LG | 2026-04-10

Loading related sources…