The Global Empirical NTK: Self-Referential Bias and Dimensionality of Gradient Descent Learning
DGX agentarXiv:2605.08746v1 Announce Type: new Abstract: In training a neural network with gradient descent (GD), each iteration induces a linear operator that governs first-order updates to a model's internal