Flatland: The Adventures of Gradient Descent with Large Step Sizes
arXiv:2606.06722v1 Announce Type: new Abstract: The training of neural networks often entails objective functions that are not globally L-smooth. For these functions, it is both theoretically and prac