When Both Layers Learn: Training Dynamics of Representing Linear Models via ReLU Networks
DGX agentarXiv:2606.04476v1 Announce Type: new Abstract: In this paper, we study the gradient descent dynamics for jointly training both layers of a one-hidden-layer ReLU network to fit a linear target functio