Research
Geometric Layer-wise Approximation Rates for Deep Networks
arXiv:2604.20219v1 Announce Type: new Abstract: Depth is widely viewed as a central contributor to the success of deep neural networks, whereas standard neural network approximation theory typically p
arXiv:2604.20219v1 Announce Type: new Abstract: Depth is widely viewed as a central contributor to the success of deep neural networks, whereas standard neural network approximation theory typically provides guarantees only for the final output and leaves the role of intermediate layers largely unclear. We address this gap by developing a quantitative framework in which depth admits a precise scale-dependent interpretation. Specifically, we design a single shared mixed-activation architecture of fixed width 2dN+d+2 and any prescribed finite depth such that each intermediate readout Phi_ell is itself an approximant to the target function f. For fin L^p([0,1]^d) with pin [1,infty), the approximation error of Phi_ell is controlled by (2d+1) times the L^p modulus of continuity at the geometric scale N^{-ell} for all ell. The estimate reduces to the geometric rate (2d+1)N^{-ell} if f is 1-Lipschitz. Our network design is inspired by multigrade deep learning, where depth serves as a progressive refinement mechanism: each new correction targets residual information at a finer scale while the earlier correction terms remain part of the later readouts, yielding a nested architecture that supports adaptive refinement without redesigning the preceding network.
Related
- Quantitative Approximation Rates for Group Equivariant Learning
- Collective Kernel EFT for Pre-activation ResNets
- Time-Frequency Analysis for Neural Networks
- Stochastic Gradient Descent in the Saddle-to-Saddle Regime of Deep Linear Networks
Source: arXiv cs.LG | 2026-04-23