Research
Learning from the Descent Direction: Adaptive Gradient Descent under One-Sided Holder Regularity
arXiv:2607.22906v1 Announce Type: new Abstract: We study adaptive gradient descent for continuously differentiable, possibly nonconvex objectives under one-sided Holder regularity. Unlike classical Ho
arXiv:2607.22906v1 Announce Type: new Abstract: We study adaptive gradient descent for continuously differentiable, possibly nonconvex objectives under one-sided Holder regularity. Unlike classical Holder- or Lipschitz-gradient assumptions, which control the full gradient variation, our condition bounds only the directional term appearing in the descent inequality. This can allow less conservative step sizes when large gradient changes are orthogonal to, or favorable along, the update direction. We propose an adaptive scalar-step method based on an estimate of positive one-sided Holder curvature, combined with a simple sufficient-decrease safeguard. For nonconvex objectives on a convex region containing the accepted update segments, we prove an explicit best-iterate stationarity bound with a rate determined by the Holder exponent. Unlike predetermined diminishing step-size schemes, the method adapts to the local descent geometry. We evaluate the approach on two full-batch benchmarks designed to separate directional curvature from full gradient variation. On a binary classification problem, the method achieves the lowest final cross-entropy, objective value, and gradient norm, together with the largest classification margin among the compared scalar gradient methods. On a nonconvex Holder regression problem, it attains the lowest final objective gap and gradient norm. These results indicate that one-sided Holder curvature is an effective adaptive step-size signal when full-gradient variation is inflated by directions that do not hinder descent.
Related
- A short proof of near-linear convergence of adaptive gradient descent under fourth-order growth and convexity
- Decoupling Variance and Scale-Invariant Updates in Adaptive Gradient Descent for Unified Vector and Matrix Optimization
- Learning rate adaptive stochastic gradient descent optimization methods: numerical simulations for deep learning methods for partial differential equations and convergence analyses
- On the Interaction of Batch Noise, Adaptivity, and Compression, under (L_0,L_1)-Smoothness: An SDE Approach
- Fisher-Geometric Diffusion in Stochastic Gradient Descent: Optimal Rates, Oracle Complexity, and Information-Theoretic Limits
Source: arXiv cs.LG | 2026-07-28