Learning Dynamics Reveal a Hierarchy of Weight-Induced Layerwise Gram Metrics
arXiv:2606.09744v1 Announce Type: new Abstract: We study feed-forward ReLU networks with fixed readout and quadratic loss. The aim is to rewrite gradient descent not primarily as a dynamics in weight