Model Releases

A law of robustness for two-layer neural networks with arbitrary weights

arXiv:2607.07778v1 Announce Type: new Abstract: Bubeck, Li and Nagaraj conjectured that, for generic data, any two-layer neural network with m neurons that fits n noisy labels must have Lipschitz cons

DGX agentpaper
model-releasesarxiv-cs-lg

arXiv:2607.07778v1 Announce Type: new Abstract: Bubeck, Li and Nagaraj conjectured that, for generic data, any two-layer neural network with m neurons that fits n noisy labels must have Lipschitz constant at least of order sqrt{n/m}, with no restriction on the size of the weights. Bubeck and Sellke proved a universal version of this law for Lipschitz-parameterized classes, but under a polynomial bound on the parameters; at depth three that boundedness hypothesis is genuinely necessary. The two-layer unbounded-weight case requires a different argument. We prove the conjectured law, up to one logarithmic factor, for every continuous piecewise-linear activation, in particular for ReLU networks. For data drawn uniformly from S^{d-1}, dge3, or from N(0,I_d/d), labels in [-1,1] with noise level sigma^2>0, and any width-m two-layer network with arbitrary real weights, biases and affine skip connection, fitting the data arepsilon below the noise floor forces Lip(f)ge c,arepsilonsqrt{n/(ar mlog(Car m nd/arepsilon))}, ar m=(K-1)m+1, with high probability. A realized-kink-count version holds on the same event: every realized two-layer piecewise-linear function with k(f)le n distinct kink hyperplanes obeys the bound with ar m replaced by k(f)+1, irrespective of how many redundant hidden units parameterize it. The proof replaces parameter-space covering, impossible for unbounded weights, by a function-space covering. The central deterministic ingredient is a rigidity lemma: on B_2, and on S^{d-1} for dge3, the coefficient of each canonical kink is controlled by the Lipschitz constant of the realized function, because kinks on distinct hyperplanes cannot cancel at generic points. Rigidity genuinely fails at d=2, and an explicit two-layer ReLU interpolant with O(1) Lipschitz constant at width 2n matches the law at the overparameterized endpoint.

Source: arXiv cs.LG | 2026-07-10

Loading related sources…