Fetching the paper…
Reading the bibliography…
Recent findings (e.g., arXiv:2103.00065) demonstrate that modern neural networks trained by full-batch gradient descent typically enter a regime called Edge of Stability (EOS).
Pathological spectra of the fisher information metric and its variants in deep neural networks
Ryo Karakida, Shotaro Akaho, and Shun-ichi Amari · 1910
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Earlier work this paper cites.
Second-order optimization for neural networks
James Martens · 2016
Earlier work this paper cites.
Eigenvalues of the hessian in deep learning: Singularity and beyond
Levent Sagun, Leon Bottou, and Yann LeCun · 2016
Earlier work this paper cites.
Empirical analysis of the hessian of over-parametrized neural networks
Levent Sagun, Utku Evci, V Ugur Guney, Yann Dauphin, and Leon Bottou · 2017
Earlier work this paper cites.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2018
Earlier work this paper cites.
Gradient descent provably optimizes over-parameterized neural networks
Simon S Du, Xiyu Zhai, Barnabas Poczos, and Aarti Singh · 2018
Earlier work this paper cites.
pytorchhessian-eigentings: efficient pytorch hessian eigendecomposition, october 2018
Noah Golmant, Zhewei Yao, Amir Gholami, Michael Mahoney, and Joseph Gonzalez · 2018
Earlier work this paper cites.
Neural tangent kernel: Convergence and generalization in neural networks
Arthur Jacot, Franck Gabriel, and Clément Hongler · 2018
Earlier work this paper cites.
The full spectrum of deepnet hessians at scale: Dynamics with sgd training and sample size
Vardan Papyan · 2018
Earlier work this paper cites.
How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective
Lei Wu, Chao Ma, et al · 2018
Cited alongside, same era.
Chen Xing, Devansh Arpit, Christos Tsirigotis, and Yoshua Bengio · 2018
Cited alongside, same era.
Fine-grained analysis of optimization and generalization for overparameterized two-layer neural networks
Sanjeev Arora, Simon Du, Wei Hu, Zhiyuan Li, and Ruosong Wang · 2019
Cited alongside, same era.
On lazy training in differentiable programming
Lenaic Chizat, Edouard Oyallon, and Francis Bach · 2019
Cited alongside, same era.
Gradient descent finds global minima of deep neural networks
Simon Du, Jason Lee, Haochuan Li, Liwei Wang, and Xiyu Zhai · 2019
Cited alongside, same era.
The break-even point on optimization trajectories of deep neural networks
Stanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit, Jacek Tabor, Kyunghyun Cho, and Krzysztof Geras · 2020
Later among the works it cites.
The large learning rate phase of deep learning: the catapult mechanism
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2020
Later among the works it cites.
Dissecting hessian: Understanding common structure of hessian in neural networks
Yikai Wu, Xingyu Zhu, Chenwei Wu, Annie Wang, and Rong Ge · 2020
Later among the works it cites.
Gradient descent on neural networks typically occurs at the edge of stability
Jeremy M Cohen, Simran Kaur, Yuanzhi Li, J Zico Kolter, and Ameet Talwalkar · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Cited alongside, same era.
Vardan Papyan · 2019
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Cited alongside, same era.
Sharpness-aware minimization for efficiently improving generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2020
Cited alongside, same era.
The surprising simplicity of the early-time learning dynamics of neural networks
Wei Hu, Lechao Xiao, Ben Adlam, and Jeffrey Pennington · 2020
Cited alongside, same era.
The normalization method for alleviating pathological sharpness in wide neural networks
Ryo Karakida, Shotaro Akaho, and Shun-ichi Amari
Cited in the paper.
Zhiyuan Li, Tianhao Wang, and Sanjeev Arora · 2021
Later among the works it cites.
The implicit bias of minima stability: A view from function space
Rotem Mulayoff, Tomer Michaeli, and Daniel Soudry · 2021
Later among the works it cites.
Understanding the unstable convergence of gradient descent
Kwangjun Ahn, Jingzhao Zhang, and Suvrit Sra · 2022
Closest in time.
Understanding gradient descent on edge of stability in deep learning
Sanjeev Arora, Zhiyuan Li, and Abhishek Panigrahi · 2022
Closest in time.
Understanding the generalization benefit of normalization layers: Sharpness reduction
Kaifeng Lyu, Zhiyuan Li, and Sanjeev Arora · 2022
Closest in time.
The multiscale structure of neural network loss functions: The effect on optimization and origin
Chao Ma, Lei Wu, and Lexing Ying · 2022
Closest in time.