Fetching the paper…
Reading the bibliography…
The vast majority of successful deep neural networks are trained using variants of stochastic gradient descent (SGD) algorithms.
Some methods of speeding up the convergence of iteration methods
Boris T Polyak · 1964
Earlier work this paper cites.
Iterative procedures for nonlinear integral equations
Donald G Anderson · 1965
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate o (1/kˆ 2)
Yurii E Nesterov · 1983
Earlier work this paper cites.
Using fast weights to deblur old memories
Geoffrey E Hinton and David C Plaut · 1987
Earlier work this paper cites.
Efficient estimations from a slowly convergent robbins-monro process
David Ruppert · 1988
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Boris T Polyak and Anatoli B Juditsky · 1992
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Mitchell P Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini · 1993
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Sharpness in rates of convergence for cg and symmetric lanczos methods
Ren-cang Li · 2005
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Neural networks: tricks of the trade , volume 7700
Grégoire Montavon, Geneviève Orr, and Klaus-Robert Müller · 2012
Earlier work this paper cites.
Extrapolation methods: theory and practice , volume 2
Claude Brezinski and M Redivo Zaglia · 2013
Earlier work this paper cites.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Earlier work this paper cites.
No more pesky learning rates
Tom Schaul, Sixin Zhang, and Yann LeCun · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Cited alongside, same era.
On using very large target vocabulary for neural machine translation
Sébastien Jean, Kyunghyun Cho, Roland Memisevic, and Yoshua Bengio · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
New insights and perspectives on the natural gradient method
James Martens · 2014
Cited alongside, same era.
Regularizing and optimizing lstm language models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Later among the works it cites.
On the ineffectiveness of variance reduced optimization for deep learning, 2018
Aaron Defazio and Léon Bottou · 2018
Later among the works it cites.
Loss surfaces, mode connectivity, and fast ensembling of dnns
Timur Garipov, Pavel Izmailov, Dmitrii Podoprikhin, Dmitry P Vetrov, and Andrew Gordon Wilson · 2018
Later among the works it cites.
Averaging weights leads to wider optima and better generalization
Pavel Izmailov, Dmitrii Podoprikhin, Timur Garipov, Dmitry Vetrov, and Andrew Gordon Wilson · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Cited alongside, same era.
Adaptive restart for accelerated gradient schemes
Brendan O’Donoghue and Emmanuel Candes · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Analysis and design of optimization algorithms via integral quadratic constraints
Laurent Lessard, Benjamin Recht, and Andrew Packard · 2016
Cited alongside, same era.
Wide residual networks
Sergey Zagoruyko and Nikos Komodakis · 2016
Cited alongside, same era.
Katyusha: The first direct acceleration of stochastic gradient methods
Zeyuan Allen-Zhu · 2017
Cited alongside, same era.
Improved regularization of convolutional neural networks with cutout
Terrance DeVries and Graham W Taylor · 2017
Cited alongside, same era.
Later among the works it cites.
Xianyan Jia, Shutao Song, Wei He, Yangzihao Wang, Haidong Rong, Feihu Zhou, Liqiang Xie, Zhenyu Guo, Yuanzhou Yang, Liwei Yu, et al · 2018
Later among the works it cites.
Aggregated momentum: Stability through passive damping
James Lucas, Richard Zemel, and Roger Grosse · 2018
Later among the works it cites.
Reptile: a scalable metalearning algorithm
Alex Nichol, Joshua Achiam, and John Schulman · 2018
Later among the works it cites.
Second-order optimization method for large mini-batch: Training resnet-50 on imagenet in 35 epochs, 2018
Kazuki Osawa, Yohei Tsuji, Yuichiro Ueno, Akira Naruse, Rio Yokota, and Satoshi Matsuoka · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Later among the works it cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern · 2018
Later among the works it cites.
Understanding short-horizon bias in stochastic meta-optimization
Yuhuai Wu, Mengye Ren, Renjie Liao, and Roger Grosse · 2018
Later among the works it cites.
The unusual effectiveness of averaging in gan training
Stefan Winkler Kim-Hui Yap Georgios Piliouras Vijay Chandrasekhar Yasin Yazıcı, Chuan-Sheng Foo · 2018
Later among the works it cites.
Imagenet training in minutes
Yang You, Zhao Zhang, Cho-Jui Hsieh, James Demmel, and Kurt Keutzer · 2018
Later among the works it cites.
Which algorithmic choices matter at which batch sizes? insights from a noisy quadratic model
Guodong Zhang, Lala Li, Zachary Nado, James Martens, Sushant Sachdeva, George E Dahl, Christopher J Shallue, and Roger Grosse · 2019
Closest in time.