Safety
Beyond Pessimism: Offline Learning in KL-regularized Games
arXiv:2604.06738v1 Announce Type: cross Abstract: We study offline learning in KL-regularized two-player zero-sum games, where policies are optimized under a KL constraint to a fixed reference policy.
arXiv:2604.06738v1 Announce Type: cross Abstract: We study offline learning in KL-regularized two-player zero-sum games, where policies are optimized under a KL constraint to a fixed reference policy. Prior work relies on pessimistic value estimation to handle distribution shift, yielding only widetilde{O}(1/sqrt n) statistical rates. We develop a new pessimism-free algorithm and analytical framework for KL-regularized games, built on the smoothness of KL-regularized best responses and a stability property of the Nash equilibrium induced by skew symmetry. This yields the first widetilde{O}(1/n) sample complexity bound for offline learning in KL-regularized zero-sum games, achieved entirely without pessimism. We further propose an efficient self-play policy optimization algorithm and prove that, with a number of iterations linear in the sample size, it achieves the same fast widetilde{O}(1/n) statistical rate as the minimax estimator.
Related
- Adaptive Replay Buffer for Offline-to-Online Reinforcement Learning
- Multi-agent Reach-avoid MDP via Potential Games and Low-rank Policy Structure
- MDP modeling for multi-stage stochastic programs
- GIFT: Group-Relative Implicit Fine-Tuning Integrates GRPO with DPO and UNA
Source: arXiv cs.LG | 2026-04-10