Fetching the paper…
Reading the bibliography…
We derive a simple and model-independent formula for the change in the generalization gap due to a gradient descent update.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Flat minima
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
On-line learning for very large data sets
Léon Bottou and Yann Le Cun · 2005
Earlier work this paper cites.
The tradeoffs of large scale learning
Léon Bottou and Olivier Bousquet · 2008
Earlier work this paper cites.
Path-sgd: Path-normalized optimization in deep neural networks
Behnam Neyshabur, Ruslan R Salakhutdinov, and Nati Srebro · 2015
Earlier work this paper cites.
Adding gradient noise improves learning for very deep networks
Arvind Neelakantan, Luke Vilnis, Quoc V Le, Ilya Sutskever, Lukasz Kaiser, Karol Kurach, and James Martens · 2015
Earlier work this paper cites.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2016
Earlier work this paper cites.
Understanding deep learning requires rethinking generalization
Chiyuan Zhang, Samy Bengio, Moritz Hardt, Benjamin Recht, and Oriol Vinyals · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang · 2016
Earlier work this paper cites.
Entropy-sgd: Biasing gradient descent into wide valleys
Pratik Chaudhari, Anna Choromanska, Stefano Soatto, Yann LeCun, Carlo Baldassi, Christian Borgs, Jennifer Chayes, Levent Sagun, and Riccardo Zecchina · 2016
Cited alongside, same era.
Geometry of optimization and implicit regularization in deep learning
Behnam Neyshabur, Ryota Tomioka, Ruslan Salakhutdinov, and Nathan Srebro · 2017
Cited alongside, same era.
Three factors influencing minima in sgd
Stanisław Jastrzębski, Zachary Kenton, Devansh Arpit, Nicolas Ballas, Asja Fischer, Yoshua Bengio, and Amos Storkey · 2017
Cited alongside, same era.
On the diffusion approximation of nonconvex stochastic gradient descent
Wenqing Hu, Chris Junchi Li, Lei Li, and Jian-Guo Liu · 2017
Cited alongside, same era.
SGD Implicitly Regularizes Generalization Error
Daniel A. Roberts · 2018
Later among the works it cites.
A bayesian perspective on generalization and stochastic gradient descent
Samuel L. Smith and Quoc V. Le · 2018
Later among the works it cites.
An alternative view: When does sgd escape local minima?
Robert Kleinberg, Yuanzhi Li, and Yang Yuan · 2018
Later among the works it cites.
The regularization effects of anisotropic noise in stochastic gradient descent
Zhanxing Zhu, Jingfeng Wu, Lei Wu, Jinwen Ma, and Bing Yu · 2018
Later among the works it cites.
Energy–entropy competition and the effectiveness of stochastic gradient descent in machine learning
Yao Zhang, Andrew M. Saxe, Madhu S. Advani, and Alpha A. Lee · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Laurent Dinh, Razvan Pascanu, Samy Bengio, and Yoshua Bengio · 2017
Cited alongside, same era.
Stochastic gradient descent as approximate bayesian inference
Stephan Mandt, Matthew D Hoffman, and David M Blei · 2017
Cited alongside, same era.
Pratik Chaudhari and Stefano Soatto · 2017
Cited alongside, same era.
Gradient diversity: a key ingredient for scalable distributed learning
Dong Yin, Ashwin Pananjady, Max Lam, Dimitris Papailiopoulos, Kannan Ramchandran, and Peter Bartlett · 2017
Cited alongside, same era.
Andrew K Lampinen and Surya Ganguli · 2018
Later among the works it cites.
Fluctuation-dissipation relations for stochastic gradient descent
Sho Yaida · 2018
Later among the works it cites.
On the origin of implicit regularization in stochastic gradient descent
Samuel L Smith, Benoit Dherin, David Barrett, and Soham De · 2021
Closest in time.