Explaining Near-Zero Hessian Eigenvalues Through Approximate Symmetries in Neural Networks
DGX agentarXiv:2607.07845v1 Announce Type: new Abstract: The Hessian of the training loss governs the local geometry of the loss landscape, yet despite existing explanations for its largest eigenvalues, the or