Fetching the paper…
Reading the bibliography…
Most existing analyses of (stochastic) gradient descent rely on the condition that for $L$-smooth costs, the step size is less than $2/L$.
Global stability of dynamical systems
M. Shub · 2013
Earlier work this paper cites.
Gradient descent only converges to minimizers
J. D. Lee, M. Simchowitz, M. I. Jordan, and B. Recht · 2016
Earlier work this paper cites.
Three factors influencing minima in SGD
S. Jastrzkebski, Z. Kenton, D. Arpit, N. Ballas, A. Fischer, Y. Bengio, and A. Storkey · 2017
Earlier work this paper cites.
Gradient descent only converges to minimizers: Non-isolated critical points and invariant regions
I. Panageas and G. Piliouras · 2017
Earlier work this paper cites.
Neural tangent kernel: convergence and generalization in neural networks
A. Jacot, F. Gabriel, and C. Hongler · 2018
Earlier work this paper cites.
On the relation between the sharpest directions of dnn loss and the sgd step length
S. Jastrzkebski, Z. Kenton, N. Ballas, A. Fischer, Y. Bengio, and A. Storkey · 2018
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Y. Li and Y. Liang · 2018
Cited alongside, same era.
How SGD selects the global minima in over-parameterized learning: A dynamical stability perspective
L. Wu, C. Ma, et al · 2018
Cited alongside, same era.
C. Xing, D. Arpit, C. Tsirigotis, and Y. Bengio · 2018
Cited alongside, same era.
Wide neural networks of any depth evolve as linear models under gradient descent
J. Lee, L. Xiao, S. Schoenholz, Y. Bahri, R. Novak, J. Sohl-Dickstein, and J. Pennington · 2019
Cited alongside, same era.
The large learning rate phase of deep learning: the catapult mechanism
A. Lewkowycz, Y. Bahri, E. Dyer, J. Sohl-Dickstein, and G. Gur-Ari · 2020
Later among the works it cites.
Gradient descent on neural networks typically occurs at the edge of stability
J. Cohen, S. Kaur, Y. Li, J. Z. Kolter, and A. Talwalkar · 2021
Later among the works it cites.
Understanding gradient descent on edge of stability in deep learning
S. Arora, Z. Li, and A. Panigrahi · 2022
Closest in time.
The multiscale structure of neural network loss functions: The effect on optimization and origin
C. Ma, D. Kunin, L. Wu, and L. Ying · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…