Model Releases
Scaling the Memory of Balanced Adam
arXiv:2605.10119v1 Announce Type: new Abstract: Recent evidence suggests that Adam performs robustly when its momentum parameters are tied, eta_1=eta_2, reducing the optimizer to a single remaining pa
arXiv:2605.10119v1 Announce Type: new Abstract: Recent evidence suggests that Adam performs robustly when its momentum parameters are tied, eta_1=eta_2, reducing the optimizer to a single remaining parameter. However, the value of this parameter is still poorly understood. We argue that, in balanced Adam, eta should not be treated as a dimensionless constant: it defines a statistical memory horizon H_eta=(1-eta)^{-1}. In terms of the effective learning horizon T_{ES}, estimated from the validation trajectory, we study the refresh count R_eta=(1-eta)T_{ES}, which measures how many times Adam renews its internal statistics during the useful phase of training. Across 11 vision and language experiments, we find that choosing eta so that R_etaapprox1000 selects different beta values depending on the training scale, yet improves robustness over the best fixed-beta baseline. Compared with the strongest fixed choice eta=0.94377, the refresh rule improves worst-case robustness, reducing the global maximum validation gap by 33.4%, while bringing all 11 runs within 1% of their validation oracle. These results suggest that the remaining hyperparameter of balanced Adam is better understood as a memory-scale variable than as a fixed constant. This provides a simple budget-aware perspective on optimizer scaling and opens a path toward treating Adam's momentum as part of the learning dynamics rather than as a static default.
Source: arXiv cs.LG | 2026-05-12