Safety

Unregularized Convergence of Single-Loop, Entropy-Regularized Natural Actor-Critic

arXiv:2608.19587v1 Announce Type: new Abstract: While entropy regularization is widely used to stabilize and accelerate Natural Policy Gradient methods, its ability to yield faster convergence rates f

DGX agentpaper
safetyarxiv-cs-lg

arXiv:2608.19587v1 Announce Type: new Abstract: While entropy regularization is widely used to stabilize and accelerate Natural Policy Gradient methods, its ability to yield faster convergence rates for the unregularized objective remains underexplored. Existing analyses often rely on double-loop architectures and invoke a linear entropy penalty. To bridge the gap between theory and practice, we analyze a single-loop, entropy-regularized Natural Actor-Critic algorithm under compatible linear function approximation. By training an uncentered critic, our critic tracking can remain stable even as the training policy approaches determinism and the Fisher information matrix degenerates. We focus on two primary regimes for the optimization landscape: a Stochastic Regime, where we fuse coupled actor-critic updates into a joint Lyapunov recurrence, and a Deterministic Regime, where we pivot to a Policy Mirror Descent framework to circumvent the collapse of Euclidean geometry. By exploiting a positive Minimal Action Gap in the unregularized Markov decision process, we introduce an Exponential Translation mechanism that maps the regularized gap to the unregularized one up to an exponentially decaying tail. By tuning the fixed temperature, our algorithm achieves accelerated unregularized convergence rates, up to approximation-error terms: ilde{O}(T_{total}^{-1}) in the Stochastic Regime, and ilde{O}(T_{total}^{-2/3}) for the average iterate alongside ilde{O}(T_{total}^{-1/3}) for the last iterate in the Deterministic Regime. Here, T_{total} denotes the total number of stochastic critic updates (or Monte Carlo rollouts). Furthermore, in the tabular setting, our positive-action-gap analysis yields a ilde{O}(T_{total}^{-2/3}) average-iterate rate, surpassing the O(T_{total}^{-1/2}) worst-case statistical barrier that applies without a positive action margin.

Related

Source: arXiv cs.LG | 2026-08-21

Loading related sources…