Research
Error Bound Analysis for the Regularized Loss of Deep Linear Neural Networks
arXiv:2502.11152v4 Announce Type: replace-cross Abstract: The optimization foundations of deep linear networks have recently received significant attention. However, due to their inherent non-convexit
arXiv:2502.11152v4 Announce Type: replace-cross Abstract: The optimization foundations of deep linear networks have recently received significant attention. However, due to their inherent non-convexity and hierarchical structure, analyzing the loss functions of deep linear networks remains a challenging task. In this work, we study the local geometry of the regularized squared loss of deep linear networks around each critical point. Specifically, we obtain a closed-form characterization of the critical point set building on existing results and establish an error bound for the regularized loss under mild conditions on network width and regularization parameters. Notably, this error bound quantifies the distance from a point to the critical point set in terms of the current gradient norm, which can be used to derive linear convergence of first-order methods. To support our theoretical findings, we conduct numerical experiments and demonstrate that gradient descent converges linearly to a critical point when optimizing the regularized loss of deep linear networks.
Source: arXiv cs.LG | 2026-08-11