Model Releases
Why Adam Can Beat SGD: Second-Moment Normalization Yields Sharper Tails
arXiv:2603.03099v5 Announce Type: replace-cross Abstract: Despite Adam demonstrating faster empirical convergence than SGD in many applications, much of the existing theory yields guarantees essential
arXiv:2603.03099v5 Announce Type: replace-cross Abstract: Despite Adam demonstrating faster empirical convergence than SGD in many applications, much of the existing theory yields guarantees essentially comparable to those of SGD, leaving the empirical performance gap insufficiently explained. In this paper, we uncover a key second-moment normalization in Adam and develop a stopping-time/martingale analysis that provably distinguishes Adam from SGD under the classical bounded variance model (a second moment assumption). In particular, we establish the first theoretical separation between the high-probability convergence behaviors of the two methods: Adam achieves a elta^{-1/2} dependence on the confidence parameter elta, whereas corresponding high-probability guarantee for SGD necessarily incurs at least a elta^{-1} dependence.
Related
- AgentOpt v0.1 Technical Report: Client-Side Optimization for LLM-Based Agent
- VOLTA: The Surprising Ineffectiveness of Auxiliary Losses for Calibrated Deep Learning
- JoyAI-LLM Flash: Advancing Mid-Scale LLMs with Token Efficiency
- ODYN: An All-Shifted Non-Interior-Point Method for Quadratic Programming in Robotics and AI
Source: arXiv cs.AI | 2026-04-13