Research

Optimal Training-Time Scaling in Gradual Adaptation

arXiv:2608.04927v1 Announce Type: new Abstract: In gradual adaptation, how should the training time on each task change as the number of intermediate tasks increases? We study this question for overpa

DGX agentpaper
researcharxiv-cs-lg

arXiv:2608.04927v1 Announce Type: new Abstract: In gradual adaptation, how should the training time on each task change as the number of intermediate tasks increases? We study this question for overparameterized linear regression tasks that change smoothly and share a zero-loss solution. With N tasks and training time s_N on each, the final learning progress converges to a continuum curve when Ns_Noau. The limiting progress is Theta(au) for small au and Theta(au^{-1}) for large au, so both very short and very long training produce little progress. It follows that optimal per-task training times scale as s_N^star=Theta(N^{-1}), equivalently Ns_N^star=Theta(1). Experiments on gradually rotated MNIST and a natural Yearbook time shift are consistent with less per-task training as the path is divided more finely.

Source: arXiv cs.LG | 2026-08-06

Loading related sources…