Fetching the paper…
Reading the bibliography…
We show that stochastic gradient descent (SGD) escapes from sharp minima exponentially fast even before SGD reaches stationary distribution.
The activated complex in chemical reactions
Henry Eyring · 1935
Earlier work this paper cites.
Brownian motion in a field of force and the diffusion model of chemical reactions
H A Kramers · 1940
Earlier work this paper cites.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Escape from a metastable state
Peter Hanggi · 1986
Earlier work this paper cites.
Training a 3-node neural network is np-complete
Avrim L Blum and Ronald L Rivest · 1992
Earlier work this paper cites.
Simplifying neural nets by discovering flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1995
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Ordinary differential equations and dynamical systems
Gerald Teschl · 2000
Earlier work this paper cites.
On the stable equilibrium points of gradient systems
P-A Absil and K Kurdyka · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Large Deviations Techniques and Applications
Amir Dembo and Ofer Zeitouni · 2010
Earlier work this paper cites.
Stopped diffusion processes: Boundary corrections and overshoot
Emmanuel Gobet and Stéphane Menozzi · 2010
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Léon Bottou · 2010
Earlier work this paper cites.
A method for scribe distinction in medieval manuscripts using page layout features
Claudio De Stefano, Francesco Fontanella, Marilena Maniaci, and Alessandra Scotto di Freca · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Random Perturbations of Dynamical Systems 3rd Ed
Mark I Freidlin and Alexander D Wentzell · 2012
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Cited alongside, same era.
A variational analysis of stochastic gradient algorithms
Stephan Mandt, Matthew Hoffman, and David Blei · 2016
Cited alongside, same era.
Monte-Carlo Methods and Stochastic Processes: From Linear to Non-Linear
Emmanuel Gobet · 2016
Cited alongside, same era.
Understanding deep learning requires rethinking generalization (2016)
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2017
Cited alongside, same era.
Gintare Karolina Dziugaite and Daniel M Roy · 2017
Revisiting small batch training for deep neural networks
Dominic Masters and Carlo Luschi · 2018
Later among the works it cites.
Escaping saddles with stochastic gradients
Hadi Daneshmand, Jonas Kohler, Aurelien Lucchi, and Thomas Hofmann · 2018
Later among the works it cites.
Fantastic generalization measures and where to find them
Yiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan, and Samy Bengio · 2019
Later among the works it cites.
First exit time analysis of stochastic gradient descent under heavy-tailed gradient noise
Thanh Huy Nguyen, Umut Şimşekli, Mert Gürbüzbalaban, and Gaël Richard · 2019
Later among the works it cites.
A Tail-Index analysis of stochastic gradient noise in deep neural networks
Umut Simsekli, Levent Sagun, and Mert Gurbuzbalaban · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Three factors influencing minima in sgd
Stanisław Jastrzębski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Cited alongside, same era.
Bridging the gap between constant step size stochastic gradient descent and markov chains
Aymeric Dieuleveut, Alain Durmus, and Francis Bach · 2017
Cited alongside, same era.
Non-convex learning via stochastic gradient langevin dynamics: a nonasymptotic analysis
Maxim Raginsky, Alexander Rakhlin, and Matus Telgarsky · 2017
Cited alongside, same era.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Cited alongside, same era.
On the diffusion approximation of nonconvex stochastic gradient descent
Wenqing Hu, Chris Junchi Li, Lei Li, and Jian-Guo Liu · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Cited alongside, same era.
Towards understanding generalization of deep learning: Perspective of loss landscapes
Lei Wu, Zhanxing Zhu, et al · 2017
Cited alongside, same era.
Later among the works it cites.
A scale invariant flatness measure for deep network minima
Akshay Rangamani, Nam H Nguyen, Abhishek Kumar, Dzung Phan, Sang H Chin, and Trac D Tran · 2019
Later among the works it cites.
On the heavy-tailed theory of stochastic gradient descent for deep neural networks
Umut Şimşekli, Mert Gürbüzbalaban, Thanh Huy Nguyen, Gaël Richard, and Levent Sagun · 2019
Later among the works it cites.
Non-gaussianity of stochastic gradient noise
Abhishek Panigrahi, Raghav Somani, Navin Goyal, and Praneeth Netrapalli · 2019
Later among the works it cites.
Quasi-potential as an implicit regularizer for the loss function in the stochastic gradient descent
Wenqing Hu, Zhanxing Zhu, Haoyi Xiong, and Jun Huan · 2019
Later among the works it cites.
A diffusion theory for deep learning dynamics: Stochastic gradient descent exponentially favors flat minima
Zeke Xie, Issei Sato, and Masashi Sugiyama · 2020
Later among the works it cites.
The break-even point on optimization trajectories of deep neural networks
Stanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit, Jacek Tabor, Kyunghyun Cho, and Krzysztof Geras · 2020
Later among the works it cites.
Normalized flat minima: Exploring scale invariant definition of flat minima for neural networks using pac-bayesian analysis
Yusuke Tsuzuku, Issei Sato, and Masashi Sugiyama · 2020
Later among the works it cites.
Sharpness-aware minimization for efficiently improving generalization
Pierre Foret, Ariel Kleiner, Hossein Mobahi, and Behnam Neyshabur · 2020
Later among the works it cites.
Stochastic gradient and langevin processes
Xiang Cheng, Dong Yin, Peter Bartlett, and Michael Jordan · 2020
Later among the works it cites.
Stationary behavior of constant stepsize sgd type algorithms: An asymptotic characterization
Zaiwei Chen, Shancong Mou, and Siva Theja Maguluri · 2021
Closest in time.
Minimum sharpness: Scale-invariant parameter-robustness of neural networks
Hikaru Ibayashi, Takuo Hamaguchi, and Masaaki Imaizumi · 2021
Closest in time.
Asam: Adaptive sharpness-aware minimization for scale-invariant learning of deep neural networks
Jungmin Kwon, Jeongseop Kim, Hyunseo Park, and In Kwon Choi · 2021
Closest in time.
Descending through a crowded valley-benchmarking deep learning optimizers
Robin M Schmidt, Frank Schneider, and Philipp Hennig · 2021
Closest in time.