Fetching the paper…
Reading the bibliography…
Stochastic Gradient Descent (SGD) and its variants are mainstream methods for training deep networks in practice.
The activated complex in chemical reactions
Henry Eyring · 1935
Earlier work this paper cites.
Brownian motion in a field of force and the diffusion model of chemical reactions
Hendrik Anthony Kramers · 1940
Earlier work this paper cites.
Limit distributions for sums of independent
BV Gnedenko, AN Kolmogorov, BV Gnedenko, and AN Kolmogorov · 1954
Earlier work this paper cites.
Escape from a metastable state
Peter Hanggi · 1986
Earlier work this paper cites.
Reaction-rate theory: fifty years after kramers
Peter Hänggi, Peter Talkner, and Michal Borkovec · 1990
Earlier work this paper cites.
Stochastic processes in physics and chemistry , volume 1
Nicolaas Godfried Van Kampen · 1992
Earlier work this paper cites.
Simplifying neural nets by discovering flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1995
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
The mnist database of handwritten digits
Yann LeCun · 1998
Earlier work this paper cites.
Mathematical methods for physicists, 1999
George B Arfken and Hans J Weber · 1999
Earlier work this paper cites.
In all likelihood: statistical modelling and inference using likelihood
Yudi Pawitan · 2001
Earlier work this paper cites.
Adai: Separating the effects of adaptive learning rate and momentum inertia
Zeke Xie, Xinrui Wang, Huishuai Zhang, Issei Sato, and Masashi Sugiyama · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Vector analysis and an introduction to tensor analysis
Seymour Lipschutz, Murray R Spiegel, and Dennis Spellman · 2009
Earlier work this paper cites.
Rate theories for biologists
Huan-Xiang Zhou · 2010
Earlier work this paper cites.
Bayesian learning via stochastic gradient langevin dynamics
Max Welling and Yee W Teh · 2011
Earlier work this paper cites.
Zeke Xie, Fengxiang He, Shaopeng Fu, Issei Sato, Dacheng Tao, and Masashi Sugiyama · 2011
Earlier work this paper cites.
Stable weight decay regularization
Zeke Xie, Issei Sato, and Masashi Sugiyama · 2011
Earlier work this paper cites.
The Langevin equation: with applications to stochastic problems in physics, chemistry and electrical engineering , volume 27
William Coffey and Yu P Kalmykov · 2012
Earlier work this paper cites.
Kramers’ law: Validity, derivations and generalisations
Nils Berglund · 2013
Earlier work this paper cites.
One weird trick for parallelizing convolutional neural networks
Alex Krizhevsky · 2014
Cited alongside, same era.
Approximation analysis of stochastic gradient langevin dynamics by using fokker-planck equation and ito process
Issei Sato and Hiroshi Nakagawa · 2014
Cited alongside, same era.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Cited alongside, same era.
Train faster, generalize better: Stability of stochastic gradient descent
Moritz Hardt, Ben Recht, and Yoram Singer · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
A closer look at memorization in deep networks
Devansh Arpit, Stanisław Jastrzębski, Nicolas Ballas, David Krueger, Emmanuel Bengio, Maxinder S Kanwal, Tegan Maharaj, Asja Fischer, Aaron Courville, Yoshua Bengio, et al · 2017
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
Pratik Chaudhari and Stefano Soatto · 2018
Later among the works it cites.
Reliable writer identification in medieval manuscripts through page layout features: The “avila” bible case
Claudio De Stefano, Marilena Maniaci, Francesco Fontanella, and A Scotto di Freca · 2018
Later among the works it cites.
Essentially no barriers in neural network energy landscape
Felix Draxler, Kambis Veschgini, Manfred Salmhofer, and Fred Hamprecht · 2018
Later among the works it cites.
Gradient descent happens in a tiny subspace
Guy Gur-Ari, Daniel A Roberts, and Ethan Dyer · 2018
Later among the works it cites.
An alternative view: When does sgd escape local minima?
Robert Kleinberg, Yuanzhi Li, and Yang Yuan · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2017
Cited alongside, same era.
Sharp minima can generalize for deep nets
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Cited alongside, same era.
UCI machine learning repository, 2017
Dheeru Dua and Casey Graff · 2017
Cited alongside, same era.
Gintare Karolina Dziugaite and Daniel M Roy · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks
Elad Hoffer, Itay Hubara, and Daniel Soudry · 2017
Cited alongside, same era.
Three factors influencing minima in sgd
Stanisław Jastrzębski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Cited alongside, same era.
Visualizing the loss landscape of neural nets
Hao Li, Zheng Xu, Gavin Taylor, Christoph Studer, and Tom Goldstein · 2018
Later among the works it cites.
A bayesian perspective on generalization and stochastic gradient descent
Samuel L Smith and Quoc V Le · 2018
Later among the works it cites.
Don’t decay the learning rate, increase the batch size
Samuel L Smith, Pieter-Jan Kindermans, and Quoc V Le · 2018
Later among the works it cites.
How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective
Lei Wu, Chao Ma, and E Weinan · 2018
Later among the works it cites.
Global convergence of langevin dynamics based algorithms for nonconvex optimization
Pan Xu, Jinghui Chen, Difan Zou, and Quanquan Gu · 2018
Later among the works it cites.
Hessian-based analysis of large batch training and robustness to adversaries
Zhewei Yao, Amir Gholami, Qi Lei, Kurt Keutzer, and Michael W Mahoney · 2018
Later among the works it cites.
Where is the information in a deep neural network?
Alessandro Achille and Stefano Soatto · 2019
Later among the works it cites.
On the diffusion approximation of nonconvex stochastic gradient descent
Wenqing Hu, Chris Junchi Li, Lei Li, and Jian-Guo Liu · 2019
Later among the works it cites.
First exit time analysis of stochastic gradient descent under heavy-tailed gradient noise
Thanh Huy Nguyen, Umut Simsekli, Mert Gurbuzbalaban, and Gaël Richard · 2019
Later among the works it cites.
A tail-index analysis of stochastic gradient noise in deep neural networks
Umut Simsekli, Levent Sagun, and Mert Gurbuzbalaban · 2019
Later among the works it cites.
Yusuke Tsuzuku, Issei Sato, and Masashi Sugiyama · 2019
Later among the works it cites.
Which algorithmic choices matter at which batch sizes? insights from a noisy quadratic model
Guodong Zhang, Lala Li, Zachary Nado, James Martens, Sushant Sachdeva, George Dahl, Chris Shallue, and Roger B Grosse · 2019
Later among the works it cites.
The anisotropic noise in stochastic gradient descent: Its behavior of escaping from sharp minima and regularization effects
Zhanxing Zhu, Jingfeng Wu, Bing Yu, Lei Wu, and Jinwen Ma · 2019
Later among the works it cites.
A hitting time analysis of stochastic gradient langevin dynamics
Yuchen Zhang, Percy Liang, and Moses Charikar · 2022
Closest in time.