Research
Momentum as Residual-Driven Multiplier Correction for Deep Learning Optimization
arXiv:2608.12925v1 Announce Type: new Abstract: Momentum-based optimizers are widely used in modern deep learning, yet the relations among momentum recursion, update geometry, and acceleration remain
arXiv:2608.12925v1 Announce Type: new Abstract: Momentum-based optimizers are widely used in modern deep learning, yet the relations among momentum recursion, update geometry, and acceleration remain only partially understood. We develop an extbf{A}DMM-extbf{I}nspired extbf{M}omentum (AIM) framework based on residual-penalty variable splitting, which interprets momentum as a multiplier-like correction driven by the splitting residual. AIM recovers the exponential moving average of gradients from an ADMM-style multiplier update and separates two mechanisms that are usually intertwined in practical optimizers: the residual penalty determines the update geometry, whereas the approximation of the objective-related subproblem determines the acceleration form. Building on AIM, we propose extbf{R}elativistic extbf{A}daptive gradient extbf{D}escent with extbf{A}ccelerated extbf{R}esidual (RADAR), which combines relativistic adaptive geometry, decoupled residual correction, and second-order momentum filtering to improve the update direction and momentum estimation. We establish stochastic convergence through a variance-perturbed Lyapunov drift analysis. Experiments on supervised vision learning, language modeling, and reinforcement learning show that RADAR achieves consistent improvements over strong adaptive optimizer baselines.
Related
- Adaptive Memory Momentum via a Model-Based Framework for Deep Learning Optimization
- Layerwise LQR for Geometry-Aware Optimization of Deep Networks
- DMuon: Efficient Distributed Muon Training with Near-Adam Overhead
Source: arXiv cs.LG | 2026-08-14