Fetching the paper…
Reading the bibliography…
Stochastic gradient descent (SGD) undergoes complicated multiplicative noise for the mean-square loss.
Henry Eyring, The Activated Complex in Chemical Reactions, The Journal of Chemical Physics 3
1935
Earlier work this paper cites.
Hendrik Anthony Kramers, Brownian motion in a field of force and the diffusion model of chemical reactions, Physica 7
1940
Earlier work this paper cites.
James S. Langer, Statistical theory of the decay of metastable states, Annals of Physics 54
1969
Earlier work this paper cites.
Djc MacKay, Bayesian model comparison and backprop nets, in Advances in Neural Information Processing Systems (1992)
1992
Earlier work this paper cites.
Robert S Maier and Daniel L. Stein, Escape problem for irreversible systems, Physical Review E 48
1993
Earlier work this paper cites.
Bernt Øksendal, Stochastic differential equations: an introduction with applications (Springer, Berlin, 1998)
1998
Earlier work this paper cites.
Mark I. Freidlin and Alexander D. Wentzell, Random Perturbations of Dynamical Systems (Springer, 1998)
1998
Earlier work this paper cites.
Anton Bovier, Michael Eckhoff, Véronique Gayrard, and Markus Klein, Metastability in reversible diffusion processes I. Sharp asymptotics for capacities and exit times, Journal of the European Mathematical Society 6
2004
Earlier work this paper cites.
Ronan Collobert and Jason Weston, A unified architecture for natural language processing: Deep neural networks with multitask learning, in International Conference on Machine Learning (2008)
2008
Earlier work this paper cites.
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton, Imagenet classification with deep convolutional neural networks, in Advances in Neural Information Processing Systems (2012)
2012
Earlier work this paper cites.
Geoffrey Hinton, Li Deng, Dong Yu, George E Dahl, Abdel-rahman Mohamed, Navdeep Jaitly, Andrew Senior, Vincent Vanhoucke, Patrick Nguyen, Tara N Sainath, and Others, Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups, IEEE Signal processing magazine 29
2012
Earlier work this paper cites.
2013
Earlier work this paper cites.
Issei Sato and Hiroshi Nakagawa, Approximation analysis of stochastic gradient langevin dynamics by using fokker-planck equation and ito process, in International Conference on Machine Learning (2014)
2014
Earlier work this paper cites.
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton, Deep learning, Nature 521
2015
Earlier work this paper cites.
Alessandra Bianchi and Alexandre Gaudillière, Metastable states, quasi-stationary distributions and soft measures, Stochastic Processes and their Applications 126
2016
Cited alongside, same era.
2017
Cited alongside, same era.
Nitish Shirish Keskar, Jorge Nocedal, Ping Tak Peter Tang, Dheevatsa Mudigere, and Mikhail Smelyanskiy, On large-batch training for deep learning: Generalization gap and sharp minima, in International Conference on Learning Representations (2017)
2017
Cited alongside, same era.
Elad Hoffer, Itay Hubara, and Daniel Soudry, Train longer, generalize better: closing the generalization gap in large batch training of neural networks, in Advances in Neural Information Processing Systems (2017)
2017
Cited alongside, same era.
Zhanxing Zhu, Jingfeng Wu, Bing Yu, Lei Wu, and Jinwen Ma, The anisotropic noise in stochastic gradient descent: Its behavior of escaping from sharp minima and regularization effects, in International Conference on Machine Learning (2019)
2019
Later among the works it cites.
Vardan Papyan, Measurements of three-level hierarchical structure in the outliers in the spectrum of deepnet hessians, in International Conference on Machine Learning (2019)
2019
Later among the works it cites.
2019
Later among the works it cites.
Raban Iten, Tony Metger, Henrik Wilming, Lídia Del Rio, and Renato Renner, Discovering Physical Concepts with Neural Networks, Physical Review Letters 124
2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio, Sharp minima can generalize for deep nets, in International Conference on Machine Learning (2017)
2017
Cited alongside, same era.
2017
Cited alongside, same era.
Qianxiao Li, Cheng Tai, and E. Weinan, Stochastic modified equations and adaptive stochastic gradient algorithms, in International Conference on Machine Learning (2017)
2017
Cited alongside, same era.
Yuchen Zhang, Percy Liang, and Moses Charikar, A hitting time analysis of stochastic gradient Langevin dynamics, in Proceedings of Machine Learning Research (2017)
2017
Cited alongside, same era.
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals, Understanding Deep Learning Requires Rethinking of Generalization, in International Conference on Learning Representations (2017)
2017
Cited alongside, same era.
Lei Wu, Chao Ma, and E. Weinan, How SGD selects the global minima in over-parameterized learning: A dynamical stability perspective, in Advances in Neural Information Processing Systems (2018)
2018
Cited alongside, same era.
Samuel L. Smith and Quoc V. Le, A Bayesian perspective on generalization and stochastic gradient descent, in International Conference on Learning Representations (2018)
2018
Cited alongside, same era.
2018
Cited alongside, same era.
V. Bapst, T. Keck, A. Grabska-Barwińska, C. Donner, E. D. Cubuk, S. S. Schoenholz, A. Obika, A. W.R. Nelson, T. Back, D. Hassabis, and P. Kohli, Unveiling the predictive power of static structure in glassy systems, Nature Physics 16
2020
Later among the works it cites.
2020
Later among the works it cites.
2020
Later among the works it cites.
Alireza Seif, Mohammad Hafezi, and Christopher Jarzynski, Machine learning the thermodynamic arrow of time, Nature Physics 17
2021
Closest in time.
Zeke Xie, Issei Sato, and Masashi Sugiyama, A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima, in International Conference on Learning Representations (2021)
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.