Fetching the paper…
Reading the bibliography…
Sharpness-Aware Minimization (SAM) is a highly effective regularization technique for improving the generalization of deep neural networks for various settings.
Improved sample complexities for deep networks and robust classification via an all-layer margin
Colin Wei and Tengyu Ma · 1910
Earlier work this paper cites.
The rotation of eigenvectors by a perturbation. iii
Chandler Davis and William Morton Kahan · 1970
Earlier work this paper cites.
On differentiating eigenvalues and eigenvectors
Jan R Magnus · 1985
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
A large-deviation inequality for vector-valued martingales
Thomas P Hayes · 2003
Earlier work this paper cites.
Stochastic approximation and recursive algorithms and applications , volume 35
Harold Kushner and G George Yin · 2003
Earlier work this paper cites.
Solving Ordinary Differential Equations I: Nonstiff Problems
E. Hairer, S.P. Nørsett, and G. Wanner · 2008
Earlier work this paper cites.
Stochastic approximation: a dynamical systems viewpoint , volume 48
Vivek S Borkar · 2009
Earlier work this paper cites.
A new learning algorithm for optimal stopping
Vivek S Borkar, Jervis Pinto, and Tarun Prabhu · 2009
Earlier work this paper cites.
Matrix analysis
Roger A. Horn and Charles R. Johnson · 2012
Earlier work this paper cites.
A differential equation for modeling nesterov’s accelerated gradient method: Theory and insights
Weijie Su, Stephen Boyd, and Emmanuel Candes · 2014
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Earlier work this paper cites.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Earlier work this paper cites.
Gintare Karolina Dziugaite and Daniel M Roy · 2017
Earlier work this paper cites.
Three factors influencing minima in sgd
Stanisław Jastrzebski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Earlier work this paper cites.
Stochastic modified equations and adaptive stochastic gradient algorithms
Qianxiao Li, Cheng Tai, and E Weinan · 2017
Earlier work this paper cites.
Stochastic gradient descent as approximate bayesian inference
Stephan Mandt, Matthew D Hoffman, and David M Blei · 2017
Cited alongside, same era.
Exploring generalization in deep learning
Behnam Neyshabur, Srinadh Bhojanapalli, David McAllester, and Nati Srebro · 2017
Cited alongside, same era.
The loss landscape of overparameterized neural networks
Yaim Cooper · 2018
Cited alongside, same era.
Essentially no barriers in neural network energy landscape
Felix Draxler, Kambis Veschgini, Manfred Salmhofer, and Fred Hamprecht · 2018
Cited alongside, same era.
Stochastic methods for composite and weakly convex optimization problems
John C Duchi and Feng Ruan · 2018
Cited alongside, same era.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry P Vetrov, and Andrew G Wilson · 2018
What happens after sgd reaches zero loss?–a mathematical framework
Zhiyuan Li, Tianhao Wang, and Sanjeev Arora · 2021
Later among the works it cites.
On linear stability of sgd and input-smoothness of neural networks
Chao Ma and Lexing Ying · 2021
Later among the works it cites.
Diametrical risk minimization: Theory and computations
Matthew D Norton and Johannes O Royset · 2021
Later among the works it cites.
Regularizing neural networks via adversarial model perturbation
Yaowei Zheng, Richong Zhang, and Yongyi Mao · 2021
Later among the works it cites.
Towards understanding sharpness-aware minimization
Maksym Andriushchenko and Nicolas Flammarion · 2022
Closest in time.
Understanding gradient descent on edge of stability in deep learning
Sanjeev Arora, Zhiyuan Li, and Abhishek Panigrahi · 2022
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective
Lei Wu, Chao Ma, and Weinan E · 2018
Cited alongside, same era.
Implicit regularization for deep neural networks driven by an ornstein-uhlenbeck like process
Guy Blanc, Neha Gupta, Gregory Valiant, and Paul Valiant · 2019
Cited alongside, same era.
Fantastic generalization measures and where to find them
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio · 2019
Cited alongside, same era.
Stochastic modified equations and dynamics of stochastic gradient algorithms i: Mathematical foundations
Qianxiao Li, Cheng Tai, and E Weinan · 2019
Cited alongside, same era.
Convergence rates for the stochastic gradient descent method for non-convex objective functions
Benjamin Fehrman, Benjamin Gess, and Arnulf Jentzen · 2020
Cited alongside, same era.
Sharpness-aware minimization for efficiently improving generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2020
Cited alongside, same era.
Closest in time.
Peter L Bartlett, Philip M Long, and Olivier Bousquet · 2022
Closest in time.
Adaptive gradient methods at the edge of stability
Jeremy M Cohen, Behrooz Ghorbani, Shankar Krishnan, Naman Agarwal, Sourabh Medapati, Michal Badura, Daniel Suo, David Cardoze, Zachary Nado, George E Dahl, et al · 2022
Closest in time.
Self-stabilization: The implicit bias of gradient descent at the edge of stability
Alex Damian, Eshaan Nichani, and Jason D Lee · 2022
Closest in time.
On the maximum hessian eigenvalue and generalization
Simran Kaur, Jeremy Cohen, and Zachary C Lipton · 2022
Closest in time.
Analyzing sharpness along gd trajectory: Progressive sharpening and edge of stability
Zhouzi Li, Zixuan Wang, and Jian Li · 2022
Closest in time.
Towards efficient and scalable sharpness-aware minimization
Yong Liu, Siqi Mai, Xiangning Chen, Cho-Jui Hsieh, and Yang You · 2022
Closest in time.
Understanding the generalization benefit of normalization layers: Sharpness reduction
Kaifeng Lyu, Zhiyuan Li, and Sanjeev Arora · 2022
Closest in time.
The multiscale structure of neural network loss functions: The effect on optimization and origin
Chao Ma, Lei Wu, and Lexing Ying · 2022
Closest in time.
Explicit regularization in overparametrized models via noise injection
Antonio Orvieto, Anant Raj, Hans Kersting, and Francis Bach · 2022
Closest in time.
Penalizing gradient norm for efficiently improving generalization in deep learning
Yang Zhao, Hao Zhang, and Xiuyuan Hu · 2022
Closest in time.
Surrogate gap minimization improves sharpness-aware training
Juntang Zhuang, Boqing Gong, Liangzhe Yuan, Yin Cui, Hartwig Adam, Nicha Dvornek, Sekhar Tatikonda, James Duncan, and Ting Liu · 2022
Closest in time.