Research
A unified convergence theory for adaptive first-order methods in the nonconvex case, including AdaNorm, full and diagonal AdaGrad, Shampoo and Muo
arXiv:2604.17423v1 Announce Type: new Abstract: A unified framework for first-order optimization algorithms fornonconvex unconstrained optimization is proposed that uses adaptivelypreconditioned gradi
arXiv:2604.17423v1 Announce Type: new Abstract: A unified framework for first-order optimization algorithms fornonconvex unconstrained optimization is proposed that uses adaptivelypreconditioned gradients and includes popular methods such as full anddiagonal AdaGrad, AdaNorm, as well as adpative variants of Shampoo andMuon. This framework also allows combining heterogeneous geometriesacross different groups of variables while preserving a unifiedconvergence analysis. A fully stochastic global rate-of-convergenceanalysis is conducted for all methods in the framework, with andwithout two types of momentum, using reasonable assumptions on thevariance of the gradient oracle and without assuming boundedstochastic gradients or small enough stepsize.
Related
- Trajectory-Restricted Optimization Conditions and Geometry-Aware Linear Convergence
- Adam-HNAG: A Convergent Reformulation of Adam with Accelerated Rate
- A short proof of near-linear convergence of adaptive gradient descent under fourth-order growth and convexity
- Smoothing the Edges: Smooth Optimization for Sparse Regularization using Hadamard Overparametrization
Source: arXiv cs.LG | 2026-04-21