Fetching the paper…
Reading the bibliography…
Stochastic gradient descent plays a fundamental role in nearly all applications of deep learning.
Asymmetric valleys: Beyond sharp and flat local minima
Haowei He, Gao Huang, and Yang Yuan · 1902
Earlier work this paper cites.
On the noisy gradient descent that generalizes as SGD
Jingfeng Wu, Wenqing Hu, Haoyi Xiong, Jun Huan, Vladimir Braverman, and Zhanxing Zhu · 1906
Earlier work this paper cites.
Linear mode connectivity and the lottery ticket hypothesis
Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, and Michael Carbin · 1912
Earlier work this paper cites.
A Stochastic Approximation Method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Efficient estimations from a slowly convergent robbins-monro process
David Ruppert · 1988
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
B T Polyak and A B Juditsky · 1992
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Practical recommendations for gradient-based training of deep architectures
Yoshua Bengio · 2012
Earlier work this paper cites.
No more pesky learning rates
Tom Schaul, Sixin Zhang, and Yann LeCun · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Approximation analysis of stochastic gradient langevin dynamics by using fokker-planck equation and ito process
Issei Sato and Hiroshi Nakagawa · 2014
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Cited alongside, same era.
Cyclical learning rates for training neural networks
Leslie N Smith · 2015
Cited alongside, same era.
Deep learning , volume 1
Ian Goodfellow, Yoshua Bengio, Aaron Courville, and Yoshua Bengio · 2016
Cited alongside, same era.
An investigation into neural net optimization via hessian eigenvalue density
Behrooz Ghorbani, Shankar Krishnan, and Ying Xiao · 2019
Later among the works it cites.
Gradient descent on neural networks typically occurs at the edge of stability
Jeremy M. Cohen, Simran Kaur, Yuanzhi Li, J. Zico Kolter, and Ameet Talwalkar · 2021
Later among the works it cites.
Strength of minibatch noise in SGD
Liu Ziyin, Kangqiao Liu, Takashi Mori, and Masahito Ueda · 2021
Later among the works it cites.
Git re-basin: Merging models modulo permutation symmetries, 2022
Samuel K. Ainsworth, Jonathan Hayase, and Siddhartha Srinivasa · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ilya Loshchilov and Frank Hutter · 2016
Cited alongside, same era.
Stochastic modified equations and adaptive stochastic gradient algorithms
Qianxiao Li, Cheng Tai, and Weinan E · 2017
Cited alongside, same era.
Stochastic gradient descent as approximate bayesian inference
Stephan Mandt, Matthew D Hoffman, and David M Blei · 2017
Cited alongside, same era.
Stochastic gradient descent performs variational inference, converges to limit cycles for deep networks
Pratik Chaudhari and Stefano Soatto · 2018
Cited alongside, same era.
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
Cited alongside, same era.
Approximation analysis of stochastic gradient langevin dynamics by using Fokker-Planck equation and ito process
Issei Sato and Hiroshi Nakagawa
Cited in the paper.
Juhan Bae, Paul Vicol, Jeff Z HaoChen, and Roger Grosse · 2022
Later among the works it cites.
Random initialisations performing above chance and how to find them
Frederik Benzing, Simon Schug, Robert Meier, Johannes von Oswald, Yassir Akram, Nicolas Zucchet, Laurence Aitchison, and Angelika Steger · 2022
Later among the works it cites.
Self-Stabilization: The implicit bias of gradient descent at the edge of stability
Alex Damian, Eshaan Nichani, and Jason D Lee · 2022
Later among the works it cites.
Stop wasting my time! saving days of imagenet and bert training with latest weight averaging
Jean Kaddour · 2022
Later among the works it cites.
Beyond the quadratic approximation: the multiscale structure of neural network loss landscapes
Chao Ma, Daniel Kunin, Lei Wu, and Lexing Ying · 2022
Later among the works it cites.