Fetching the paper…
Reading the bibliography…
For infinitesimal learning rates, stochastic gradient descent (SGD) follows the path of gradient flow on the full batch loss function.
Handbook of stochastic methods , volume 3
Crispin W Gardiner et al · 1985
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-Ichi Amari · 1998
Earlier work this paper cites.
Geometric numerical integration: structure-preserving algorithms for ordinary differential equations , volume 31
Ernst Hairer, Christian Lubich, and Gerhard Wanner · 2006
Earlier work this paper cites.
Stochastic gradient descent tricks
Léon Bottou · 2012
Earlier work this paper cites.
Efficient backprop
Yann A LeCun, Léon Bottou, Genevieve B Orr, and Klaus-Robert Müller · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
Wide residual networks
Sergey Zagoruyko and Nikos Komodakis · 2016
Earlier work this paper cites.
Accurate, Large Minibatch SGD: Training Imagenet in 1 Hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2017
Earlier work this paper cites.
Stochastic modified equations and adaptive stochastic gradient algorithms
Qianxiao Li, Cheng Tai, et al · 2017
Earlier work this paper cites.
Stochastic gradient descent as approximate bayesian inference
Stephan Mandt, Matthew D Hoffman, and David M Blei · 2017
Earlier work this paper cites.
Fashion-MNIST: a Novel Image Dataset for Benchmarking Machine Learning Algorithms
Han Xiao, Kashif Rasul, and Roland Vollgraf · 2017
Cited alongside, same era.
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
Pratik Chaudhari and Stefano Soatto · 2018
Cited alongside, same era.
Three Factors Influencing Minima in SGD
Stanisław Jastrzębski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2018
Cited alongside, same era.
The power of interpolation: Understanding the effectiveness of SGD in modern over-parametrized learning
Siyuan Ma, Raef Bassily, and Mikhail Belkin · 2018
Cited alongside, same era.
An empirical model of large-batch training
Sam McCandlish, Jared Kaplan, Dario Amodei, and OpenAI Dota Team · 2018
Cited alongside, same era.
The effect of network width on stochastic gradient descent and generalization: an empirical study
Daniel Park, Jascha Sohl-Dickstein, Quoc Le, and Samuel Smith · 2019
Later among the works it cites.
A tail-index analysis of stochastic gradient noise in deep neural networks
Umut Simsekli, Levent Sagun, and Mert Gurbuzbalaban · 2019
Later among the works it cites.
Fluctuation-dissipation relations for stochastic gradient descent
Sho Yaida · 2019
Later among the works it cites.
Which Algorithmic Choices Matter at Which Batch Sizes? Insights From a Noisy Quadratic Model
Guodong Zhang, Lala Li, Zachary Nado, James Martens, Sushant Sachdeva, George E. Dahl, Christopher J. Shallue, and Roger B. Grosse · 2019
Later among the works it cites.
Batch normalization biases residual blocks towards the identity function in deep networks
Soham De and Sam Smith · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
SGD implicitly regularizes generalization error
Daniel A Roberts · 2018
Cited alongside, same era.
Measuring the effects of data parallelism on neural network training
Christopher J Shallue, Jaehoon Lee, Joe Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E Dahl · 2018
Cited alongside, same era.
A bayesian perspective on generalization and stochastic gradient descent
Samuel L Smith and Quoc V Le · 2018
Cited alongside, same era.
Don’t decay the learning rate, increase the batch size
Samuel L Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V Le · 2018
Cited alongside, same era.
How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective
Lei Wu, Chao Ma, and E Weinan · 2018
Cited alongside, same era.
Towards explaining the regularization effect of initial large learning rate in training neural networks
Yuanzhi Li, Colin Wei, and Tengyu Ma · 2019
Cited alongside, same era.
Stanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit, Jacek Tabor, Kyunghyun Cho, and Krzysztof Geras · 2020
Later among the works it cites.
The large learning rate phase of deep learning: the catapult mechanism
Aitor Lewkowycz, Yasaman Bahri, Ethan Dyer, Jascha Sohl-Dickstein, and Guy Gur-Ari · 2020
Later among the works it cites.
The impact of neural network overparameterization on gradient confusion and stochastic gradient descent
Karthik A Sankararaman, Soham De, Zheng Xu, W Ronny Huang, and Tom Goldstein · 2020
Later among the works it cites.
On the generalization benefit of noise in stochastic gradient descent
Samuel L Smith, Erich Elsen, and Soham De · 2020
Later among the works it cites.
Implicit gradient regularization
David Barrett and Benoit Dherin · 2021
Closest in time.