Fetching the paper…
Reading the bibliography…
Recent studies of gradient descent with large step sizes have shown that there is often a regime with an initial increase in the largest eigenvalue of the loss Hessian (progressive sharpening), followed by a stabilization of the eigenvalue near the maximum value which allows convergence (edge of stability).
Exploring Generalization in Deep Learning
Behnam Neyshabur, Srinadh Bhojanapalli, David Mcallester, and Nati Srebro · 2017
Earlier work this paper cites.
Neural Tangent Kernel: Convergence and Generalization in Neural Networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Earlier work this paper cites.
How SGD Selects the Global Minima in Over-parameterized Learning: A Dynamical Stability Perspective
Lei Wu, Chao Ma, and Weinan E · 2018
Earlier work this paper cites.
Reconciling modern machine-learning practice and the classical bias–variance trade-off
Mikhail Belkin, Daniel Hsu, Siyuan Ma, and Soumik Mandal · 2019
Earlier work this paper cites.
An Investigation into Neural Net Optimization via Hessian Eigenvalue Density
Behrooz Ghorbani, Shankar Krishnan, and Ying Xiao · 2019
Earlier work this paper cites.
Wide Neural Networks of Any Depth Evolve as Linear Models Under Gradient Descent
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Earlier work this paper cites.
Neural Tangents: Fast and Easy Infinite Neural Networks in Python
Roman Novak, Lechao Xiao, Jiri Hron, Jaehoon Lee, Alexander A. Alemi, Jascha Sohl-Dickstein, and Samuel S. Schoenholz · 2019
Earlier work this paper cites.
The Neural Tangent Kernel in High Dimensions: Triple Descent and a Multi-Scale Theory of Generalization
Ben Adlam and Jeffrey Pennington · 2020
Earlier work this paper cites.
Beyond Linearization: On Quadratic and Higher-Order Approximation of Wide Neural Networks
Yu Bai and Jason D. Lee · 2020
Cited alongside, same era.
At Stability’s Edge: How to Adjust Hyperparameters to Preserve Minima Selection in Asynchronous Training of Neural Networks?
Niv Giladi, Mor Shpigel Nacson, Elad Hoffer, and Daniel Soudry · 2020
Cited alongside, same era.
Dynamics of Deep Neural Networks and Neural Tangent Hierarchy
Jiaoyang Huang and Horng-Tzer Yau · 2020
Cited alongside, same era.
The asymptotic spectrum of the Hessian of DNN throughout training
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2020
Cited alongside, same era.
The large learning rate phase of deep learning: The catapult mechanism
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2020
Cited alongside, same era.
Multiple Descent: Design Your Own Generalization Curve
Self-Consistent Dynamical Field Theory of Kernel Evolution in Wide Neural Networks, May 2022
Blake Bordelon and Cengiz Pehlevan · 2022
Closest in time.
Sharpness-aware Minimization for Efficiently Improving Generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2022
Closest in time.
A Loss Curvature Perspective on Training Instabilities of Deep Learning Models
Justin Gilmer, Behrooz Ghorbani, Ankush Garg, Sneha Kudugunta, Behnam Neyshabur, David Cardoze, George Edward Dahl, Zachary Nado, and Orhan Firat · 2022
Closest in time.
Analyzing Sharpness along GD Trajectory: Progressive Sharpening and Edge of Stability, July 2022
Zhouzi Li, Zixuan Wang, and Jian Li · 2022
Closest in time.
The Principles of Deep Learning Theory
Daniel A. Roberts, Sho Yaida, and Boris Hanin · 2022
Closest in time.
Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer, March 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Lin Chen, Yifei Min, Mikhail Belkin, and Amin Karbasi · 2021
Cited alongside, same era.
Greg Yang · 2021
Cited alongside, same era.
Gradient Descent on Neural Networks Typically Occurs at the Edge of Stability
Jeremy Cohen, Simran Kaur, Yuanzhi Li, J. Zico Kolter, and Ameet Talwalkar
Cited in the paper.
Adaptive Gradient Methods at the Edge of Stability, July 2022b
Jeremy M. Cohen, Behrooz Ghorbani, Shankar Krishnan, Naman Agarwal, Sourabh Medapati, Michal Badura, Daniel Suo, David Cardoze, Zachary Nado, George E. Dahl, and Justin Gilmer
Cited in the paper.
Greg Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor, Xiaodong Liu, David Farhi, Nick Ryder, Jakub Pachocki, Weizhu Chen, and Jianfeng Gao · 2022
Closest in time.
Quadratic models for understanding neural network dynamics, May 2022
Libin Zhu, Chaoyue Liu, Adityanarayanan Radhakrishnan, and Mikhail Belkin · 2022
Closest in time.