Fetching the paper…
Reading the bibliography…
Recently, researchers observed that gradient descent for deep neural networks operates in an ``edge-of-stability'' (EoS) regime: the sharpness (maximum eigenvalue of the Hessian) is often larger than stability threshold $2/\eta$ (where $\eta$ is the step size).
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Understanding the difficulty of training deep feedforward neural networks
Xavier Glorot and Yoshua Bengio · 2010
Earlier work this paper cites.
Eigenvalues of the hessian in deep learning: Singularity and beyond
Levent Sagun, Leon Bottou, and Yann LeCun · 2016
Earlier work this paper cites.
Theoretical analysis of auto rate-tuning by batch normalization
Sanjeev Arora, Zhiyuan Li, and Kaifeng Lyu · 2018
Earlier work this paper cites.
The full spectrum of deepnet hessians at scale: Dynamics with sgd training and sample size
Vardan Papyan · 2018
Earlier work this paper cites.
High-dimensional probability: An introduction with applications in data science , volume 47
Roman Vershynin · 2018
Earlier work this paper cites.
Chen Xing, Devansh Arpit, Christos Tsirigotis, and Yoshua Bengio · 2018
Earlier work this paper cites.
The break-even point on optimization trajectories of deep neural networks
Stanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit, Jacek Tabor, Kyunghyun Cho, and Krzysztof Geras · 2020
Cited alongside, same era.
The large learning rate phase of deep learning: the catapult mechanism
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2020
Cited alongside, same era.
Dissecting hessian: Understanding common structure of hessian in neural networks
Yikai Wu, Xingyu Zhu, Chenwei Wu, Annie Wang, and Rong Ge · 2020
Cited alongside, same era.
Pyhessian: Neural networks through the lens of the hessian
Zhewei Yao, Amir Gholami, Kurt Keutzer, and Michael W Mahoney · 2020
Cited alongside, same era.
Tilting the playing field: Dynamical loss functions for machine learning
Miguel Ruiz-Garcia, Ge Zhang, Samuel S Schoenholz, and Andrea J Liu · 2021
Later among the works it cites.
Understanding the unstable convergence of gradient descent
Kwangjun Ahn, Jingzhao Zhang, and Suvrit Sra · 2022
Closest in time.
Understanding gradient descent on edge of stability in deep learning
Sanjeev Arora, Zhiyuan Li, and Abhishek Panigrahi · 2022
Closest in time.
On gradient descent convergence beyond the edge of stability
Lei Chen and Joan Bruna · 2022
Closest in time.
Understanding the generalization benefit of normalization layers: Sharpness reduction
Kaifeng Lyu, Zhiyuan Li, and Sanjeev Arora · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jeremy M Cohen, Simran Kaur, Yuanzhi Li, J Zico Kolter, and Ameet Talwalkar · 2021
Cited alongside, same era.
Label noise sgd provably prefers flat global minimizers
Alex Damian, Tengyu Ma, and Jason D Lee · 2021
Cited alongside, same era.
What happens after sgd reaches zero loss?–a mathematical framework
Zhiyuan Li, Tianhao Wang, and Sanjeev Arora · 2021
Cited alongside, same era.
Robust training of neural networks using scale invariant architectures
Zhiyuan Li, Srinadh Bhojanapalli, Manzil Zaheer, Sashank Reddi, and Sanjiv Kumar
Cited in the paper.
Analyzing sharpness along gd trajectory: Progressive sharpening and edge of stability
Zhouzi Li, Zixuan Wang, and Jian Li
Cited in the paper.
Closest in time.
The multiscale structure of neural network loss functions: The effect on optimization and origin
Chao Ma, Lei Wu, and Lexing Ying · 2022
Closest in time.
Large learning rate tames homogeneity: Convergence and balancing effect
Yuqing Wang, Minshuo Chen, Tuo Zhao, and Molei Tao · 2022
Closest in time.