Fetching the paper…
Reading the bibliography…
We empirically demonstrate that full-batch gradient descent on neural network training objectives typically operates in a regime we call the Edge of Stability.
Some methods of speeding up the convergence of iteration methods
B.T. Polyak · 1964
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate o (1/kˆ 2)
Yurii E Nesterov · 1983
Earlier work this paper cites.
Properties of the momentum lms algorithm
Mehmet Ali Tugay and Yalçin Tanik · 1989
Earlier work this paper cites.
Numerical Recipes in C (2nd Ed.): The Art of Scientific Computing
William H. Press, Saul A. Teukolsky, William T. Vetterling, and Brian P. Flannery · 1992
Earlier work this paper cites.
Automatic learning rate maximization by on-line estimation of the hessian’s eigenvectors
Yann LeCun, Patrice Y. Simard, and Barak Pearlmutter · 1993
Earlier work this paper cites.
Efficient backprop
Y LeCun, L Bottou, GB Orr, and K-R Muller · 1998
Earlier work this paper cites.
Introductory lectures on convex programming volume i: Basic course
Yurii Nesterov · 1998
Earlier work this paper cites.
An introduction to difference equations
Saber Elaydi · 2005
Earlier work this paper cites.
No more pesky learning rates
Tom Schaul, Sixin Zhang, and Yann LeCun · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Earlier work this paper cites.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima, 2016
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
SECOND-ORDER OPTIMIZATION FOR NEURAL NETWORKS
James Martens · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
Stochastic variance reduction for nonconvex optimization
Sashank J Reddi, Ahmed Hefny, Suvrit Sra, Barnabas Poczos, and Alex Smola · 2016
Earlier work this paper cites.
Finding approximate local minima faster than gradient descent
Naman Agarwal, Zeyuan Allen-Zhu, Brian Bullins, Elad Hazan, and Tengyu Ma · 2017
Earlier work this paper cites.
Why momentum really works
Gabriel Goh · 2017
Earlier work this paper cites.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Earlier work this paper cites.
On the diffusion approximation of nonconvex stochastic gradient descent
Wenqing Hu, Chris Junchi Li, Lei Li, and Jian-Guo Liu · 2017
Earlier work this paper cites.
Three factors influencing minima in sgd
Stanisław Jastrzębski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Earlier work this paper cites.
Empirical analysis of the hessian of over-parametrized neural networks
Levent Sagun, Utku Evci, V Ugur Guney, Yann Dauphin, and Leon Bottou · 2017
Earlier work this paper cites.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E. Curtis, and Jorge Nocedal · 2018
Cited alongside, same era.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Cited alongside, same era.
Learning overparameterized neural networks via stochastic gradient descent on structured data
Yuanzhi Li and Yingyu Liang · 2018
Cited alongside, same era.
Step size matters in deep learning
Kamil Nar and Shankar Sastry · 2018
Cited alongside, same era.
The full spectrum of deepnet hessians at scale: Dynamics with sgd training and sample size
Vardan Papyan · 2018
Cited alongside, same era.
How does batch normalization help optimization?
Shibani Santurkar, Dimitris Tsipras, Andrew Ilyas, and Aleksander Mądry · 2018
Painless stochastic gradient: Interpolation, line-search, and convergence rates
Sharan Vaswani, Aaron Mishkin, Issam Laradji, Mark Schmidt, Gauthier Gidel, and Simon Lacoste-Julien · 2019
Later among the works it cites.
Adagrad stepsizes: sharp convergence over nonconvex landscapes
Rachel Ward, Xiaoxia Wu, and Leon Bottou · 2019
Later among the works it cites.
The anisotropic noise in stochastic gradient descent: Its behavior of escaping from sharp minima and regularization effects
Zhanxing Zhu, Jingfeng Wu, Bing Yu, Lei Wu, and Jinwen Ma · 2019
Later among the works it cites.
Understanding the role of momentum in non-convex optimization: Practical insights from a lyapunov analysis, 2020
Aaron Defazio · 2020
Later among the works it cites.
On the convergence of adam and adagrad
Alexandre Défossez, Léon Bottou, Francis Bach, and Nicolas Usunier · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The implicit bias of gradient descent on separable data
Daniel Soudry, Elad Hoffer, Mor Shpigel Nacson, Suriya Gunasekar, and Nathan Srebro · 2018
Cited alongside, same era.
How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective
Lei Wu, Chao Ma, and E Weinan · 2018
Cited alongside, same era.
Chen Xing, Devansh Arpit, Christos Tsirigotis, and Yoshua Bengio · 2018
Cited alongside, same era.
Adaptive methods for nonconvex optimization
Manzil Zaheer, Sashank Reddi, Devendra Sachan, Satyen Kale, and Sanjiv Kumar · 2018
Cited alongside, same era.
On the convergence of adaptive gradient methods for nonconvex optimization
Dongruo Zhou, Yiqi Tang, Ziyan Yang, Yuan Cao, and Quanquan Gu · 2018
Cited alongside, same era.
A convergence theory for deep learning via over-parameterization
Zeyuan Allen-Zhu, Yuanzhi Li, and Zhao Song · 2019
Cited alongside, same era.
At stability’s edge: How to adjust hyperparameters to preserve minima selection in asynchronous training of neural networks?
Niv Giladi, Mor Shpigel Nacson, Elad Hoffer, and Daniel Soudry · 2020
Later among the works it cites.
Evaluation of neural architectures trained with square loss vs cross-entropy in classification tasks, 2020
Like Hui and Mikhail Belkin · 2020
Later among the works it cites.
The asymptotic spectrum of the hessian of dnn throughout training
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2020
Later among the works it cites.
The break-even point on optimization trajectories of deep neural networks
Stanisław Jastrzębski, Maciej Szymczak, Stanislav Fort, Devansh Arpit, Jacek Tabor, Kyunghyun Cho*, and Krzysztof Geras* · 2020
Later among the works it cites.
Neural spectrum alignment: Empirical study
Dmitry Kopitkov and Vadim Indelman · 2020
Later among the works it cites.
The large learning rate phase of deep learning: the catapult mechanism
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2020
Later among the works it cites.
Hessian based analysis of sgd for deep nets: Dynamics and generalization
Xinyan Li, Qilong Gu, Yingxue Zhou, Tiancong Chen, and Arindam Banerjee · 2020
Later among the works it cites.
Toward a theory of optimization for over-parameterized systems of non-linear equations: the lessons of deep learning, 2020
Chaoyue Liu, Libin Zhu, and Mikhail Belkin · 2020
Later among the works it cites.
Unique properties of flat minima in deep networks, 2020
Rotem Mulayoff and Tomer Michaeli · 2020
Later among the works it cites.
Traces of class/cross-class structure pervade deep learning spectra, 2020
Vardan Papyan · 2020
Later among the works it cites.
Linear convergence of adaptive stochastic gradient descent
Yuege Xie, Xiaoxia Wu, and Rachel Ward · 2020
Later among the works it cites.
Large batch optimization for deep learning: Training bert in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh · 2020
Later among the works it cites.
Why gradient clipping accelerates training: A theoretical justification for adaptivity
Jingzhao Zhang, Tianxing He, Suvrit Sra, and Ali Jadbabaie · 2020
Later among the works it cites.
Implicit gradient regularization
David Barrett and Benoit Dherin · 2021
Closest in time.
Adaptive federated optimization
Sashank J. Reddi, Zachary Charles, Manzil Zaheer, Zachary Garrett, Keith Rush, Jakub Konečný, Sanjiv Kumar, and Hugh Brendan McMahan · 2021
Closest in time.
A diffusion theory for deep learning dynamics: Stochastic gradient descent exponentially favors flat minima
Zeke Xie, Issei Sato, and Masashi Sugiyama · 2021
Closest in time.