Safety
Optimizing Regret
arXiv:2607.18866v2 Announce Type: replace-cross Abstract: Building on the identity that expected regret equals the covariance between costs and decisions, this paper develops a derivative theory of th
arXiv:2607.18866v2 Announce Type: replace-cross Abstract: Building on the identity that expected regret equals the covariance between costs and decisions, this paper develops a derivative theory of the covariance regret functional. We derive the Gateaux derivative, showing that the universal steepest-descent direction is the contrarian policy -(c-ar c), while ascent yields momentum. For linear policies hatpi(c)=Ac+b, the gradient is the cost covariance matrix Sigma_c, with a zero Hessian implying boundary-optimal solutions such as the minimum-variance portfolio. We extend to constrained optimization, sign-gradient duality between regret minimization and alpha maximization, finite-sample convergence bounds paralleling Thompson Sampling, and gradient-descent algorithms requiring only input observations.
Related
- Tight Lower Bounds for the Multi-Secretary Problem via Bellman Certificates
- A Two-Timescale Primal-Dual Framework for Reinforcement Learning via Online Dual Variable Guidance
- Fast and Robust Convergence Rate for TD(0) with Linear Function Approximation, Universal Learning Steps and I.I.D. Samples
- Fitted Q Evaluation Without Bellman Completeness via Stationary Weighting
Source: arXiv cs.LG | 2026-07-31