Fetching the paper…
Reading the bibliography…
Recent work has shown that methods like SAM which either explicitly or implicitly penalize second order information can improve generalization in deep learning.
A stochastic estimator of the trace of the influence matrix for laplacian smoothing splines
Michael F Hutchinson · 1989
Earlier work this paper cites.
Training with noise is equivalent to tikhonov regularization
Chris M Bishop · 1995
Earlier work this paper cites.
The effects of adding noise during backpropagation training on a generalization performance
Guozhong An · 1996
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Learning recurrent neural networks with hessian-free optimization
James Martens and Ilya Sutskever · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Marc’aurelio Ranzato, Andrew Senior, Paul Tucker, Ke Yang, et al · 2012
Earlier work this paper cites.
Estimating or propagating gradients through stochastic neurons for conditional computation
Yoshua Bengio, Nicholas Léonard, and Aaron Courville · 2013
Earlier work this paper cites.
Revisiting natural gradient for deep networks
Razvan Pascanu and Yoshua Bengio · 2013
Earlier work this paper cites.
Optimizing Neural Networks with Kronecker-factored Approximate Curvature
James Martens and Roger Grosse · 2015
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Earlier work this paper cites.
Sergey Zagoruyko and Nikos Komodakis · 2016
Earlier work this paper cites.
Resurrecting the sigmoid in deep learning through dynamical isometry: Theory and practice
Jeffrey Pennington, Samuel Schoenholz, and Surya Ganguli · 2017
Cited alongside, same era.
Empirical analysis of the hessian of over-parametrized neural networks
Levent Sagun, Utku Evci, V Ugur Guney, Yann Dauphin, and Leon Bottou · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour, 2018
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2018
Cited alongside, same era.
Neural Tangent Kernel: Convergence and Generalization in Neural Networks
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2018
Cited alongside, same era.
On Lazy Training in Differentiable Programming
Lénaïc Chizat, Edouard Oyallon, and Francis Bach · 2019
Cited alongside, same era.
Implicit gradient regularization
David G. T. Barrett and Benoit Dherin · 2021
Later among the works it cites.
James Martens, Andy Ballard, Guillaume Desjardins, Grzegorz Swirszcz, Valentin Dalibard, Jascha Sohl-Dickstein, and Samuel S. Schoenholz · 2021
Later among the works it cites.
A deeper look at the hessian eigenspectrum of deep neural networks and its applications to regularization
Adepu Ravi Sankar, Yash Khasbage, Rahul Vigneswaran, and Vineeth N Balasubramanian · 2021
Later among the works it cites.
Analytic Insights into Structure and Rank of Neural Network Hessian Maps, July 2021
Sidak Pal Singh, Gregor Bachmann, and Thomas Hofmann · 2021
Later among the works it cites.
On the origin of implicit regularization in stochastic gradient descent, 2021
Samuel L. Smith, Benoit Dherin, David G. T. Barrett, and Soham De · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Wide Neural Networks of Any Depth Evolve as Linear Models Under Gradient Descent
Jaehoon Lee, Lechao Xiao, Samuel Schoenholz, Yasaman Bahri, Roman Novak, Jascha Sohl-Dickstein, and Jeffrey Pennington · 2019
Cited alongside, same era.
Penalizing gradient norm for efficiently improving generalization in deep learning, 2022
Yang Zhao, Hao Zhang, and Xiuyuan Hu · 2019
Cited alongside, same era.
Temperature check: Theory and practice for training models with softmax-cross-entropy losses, October 2020
Atish Agarwala, Jeffrey Pennington, Yann Dauphin, and Sam Schoenholz · 2020
Cited alongside, same era.
Sharpness-aware minimization for efficiently improving generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2020
Cited alongside, same era.
The asymptotic spectrum of the Hessian of DNN throughout training
Arthur Jacot, Franck Gabriel, and Clement Hongler · 2020
Cited alongside, same era.
New Insights and Perspectives on the Natural Gradient Method
James Martens · 2020
Cited alongside, same era.
The implicit and explicit regularization effects of dropout
Colin Wei, Sham Kakade, and Tengyu Ma · 2020
Cited alongside, same era.
Second-order regression models exhibit progressive sharpening to the edge of stability, October 2022
Atish Agarwala, Fabian Pedregosa, and Jeffrey Pennington · 2022
Later among the works it cites.
Towards understanding sharpness-aware minimization
Maksym Andriushchenko and Nicolas Flammarion · 2022
Later among the works it cites.
Sharpness-aware training for free
Jiawei Du, Zhou Daquan, Jiashi Feng, Vincent Tan, and Joey Tianyi Zhou · 2022
Later among the works it cites.
SAM operates far from home: Eigenvalue regularization as a dynamical phenomenon
Atish Agarwala and Yann Dauphin · 2023
Later among the works it cites.
Sophia: A scalable stochastic second-order optimizer for language model pre-training, 2023
Hong Liu, Zhiyuan Li, David Hall, Percy Liang, and Tengyu Ma · 2023
Later among the works it cites.
SAMBA: Regularized autoencoders perform sharpness-aware minimization
Patrik Reizinger and Ferenc Huszár · 2023
Later among the works it cites.
The Hessian perspective into the Nature of Convolutional Neural Networks
Sidak Pal Singh, Thomas Hofmann, and Bernhard Schölkopf · 2023
Later among the works it cites.