Safety
Gradient Flow Dynamics and Implicit Bias of Diagonal Linear Networks under Infinitesimal Initialization
arXiv:2607.12332v1 Announce Type: new Abstract: We study the gradient flow dynamics of diagonal linear networks for regression tasks under infinitesimal initialization. Extending Theorem 1 from Pesme
arXiv:2607.12332v1 Announce Type: new Abstract: We study the gradient flow dynamics of diagonal linear networks for regression tasks under infinitesimal initialization. Extending Theorem 1 from Pesme & Flammarion (2023), we generalize the analysis to both deep diagonal linear networks and a broader class of two-layer diagonal linear networks (as defined in Definition 4.1). Specifically, we demonstrate that the training trajectories of these models can be equivalently characterized by the proposed Algorithm 1. We further prove that this algorithm converges to the solution of a modified l_1 norm minimization problem. As a result, we establish that the implicit bias of both network architectures corresponds to a modified l_1 norm in the regime of infinitesimal initialization. Additionally, we provide insights into the underlying mechanisms governing these dynamics by identifying the Structural Invariant Manifold (SIM) (Zhao et al., 2026) as the key geometric structure that shapes the learning process.
Related
- Implicit Bias of Mirror Flow in Homogeneous Neural Networks: Sparse and Dense Feature Learning
- The Implicit Bias of Adam and Muon on Smooth Homogeneous Neural Networks
- Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputs
Source: arXiv cs.LG | 2026-07-15