Model Releases
A Statistical Difference between Single-Layer Learning and Hierarchical Learning in Wide Neural Networks
arXiv:2607.23397v1 Announce Type: new Abstract: Hierarchical neural networks are widely used in artificial intelligence, yet their mathematical properties remain incompletely understood. In the infini
arXiv:2607.23397v1 Announce Type: new Abstract: Hierarchical neural networks are widely used in artificial intelligence, yet their mathematical properties remain incompletely understood. In the infinite-width limit, two different theoretical frameworks have been proposed. One reduces deep learning to kernel regression with a fixed kernel by assuming that the parameters remain close to their initialization, whereas the other allows the parameters to move away from their initialization, requiring the kernel itself to be optimized. In this paper, we study a three-layer neural network with a finite but large number of hidden units. We show that training the input-to-hidden weights yields a smaller generalization error than keeping them fixed. Furthermore, the latter setting exhibits singularities in the parameter space, whereas the former does not. These findings indicate that singularities play an essential role even in wide neural networks.
Related
- Feature Learning in Wide Neural Networks under muP: Identifiability and Sparse-Dictionary Decomposition of the Mean-Field Limit
- Learning Sparse Compositional Functions with Norm-Constrained Neural Networks
- How does feature learning reshape the function space?
Source: arXiv cs.LG | 2026-07-28