Fetching the paper…
Reading the bibliography…
Stochastic gradient descent (SGD) with constant momentum and its variants such as Adam are the optimization algorithms of choice for training deep neural networks (DNNs).
Some methods of speeding up the convergence of iteration methods
Boris T Polyak · 1964
Earlier work this paper cites.
Convex analysis
R Tyrrell Rockafellar · 1970
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate o (1/kˆ 2)
Yurii E Nesterov · 1983
Earlier work this paper cites.
Optimal methods of smooth convex minimization
Arkaddii S Nemirovskii and Yu E Nesterov · 1985
Earlier work this paper cites.
Principal component analysis
Svante Wold, Kim Esbensen, and Paul Geladi · 1987
Earlier work this paper cites.
Introductory lectures on convex programming volume i: Basic course
Yurii Nesterov · 1998
Earlier work this paper cites.
Variational analysis and generalized differentiation I: Basic theory
Boris S Mordukhovich · 2006
Earlier work this paper cites.
Visualizing data using t-sne
Laurens van der Maaten and Geoffrey Hinton · 2008
Earlier work this paper cites.
A fast iterative shrinkage-thresholding algorithm for linear inverse problems
Amir Beck and Marc Teboulle · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Variational analysis
R Tyrrell Rockafellar and Roger J-B Wets · 2009
Earlier work this paper cites.
MNIST handwritten digit database
Yann LeCun and Corinna Cortes · 2010
Earlier work this paper cites.
Parallelized stochastic gradient descent
Martin Zinkevich, Markus Weimer, Lihong Li, and Alex J Smola · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Earlier work this paper cites.
Advances in optimizing recurrent networks
Yoshua Bengio, Nicolas Boulanger-Lewandowski, and Razvan Pascanu · 2013
Earlier work this paper cites.
Gradient methods for minimizing composite functions
Yu Nesterov · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Earlier work this paper cites.
First-order methods of smooth convex optimization with inexact oracle
Olivier Devolder, François Glineur, and Yurii Nesterov · 2014
Cited alongside, same era.
Monotonicity and restart in fast gradient methods
Pontus Giselsson and Stephen Boyd · 2014
Cited alongside, same era.
Robustness versus acceleration
Moritz Hardt · 2014
Cited alongside, same era.
Primal-dual subgradient methods for minimizing uniformly convex functions
Anatoli Iouditski and Yuri Nesterov · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
An adaptive accelerated proximal gradient method and its homotopy continuation for sparse optimization
Katyusha: The first direct acceleration of stochastic gradient methods
Zeyuan Allen-Zhu · 2017
Later among the works it cites.
Wasserstein generative adversarial networks
Martin Arjovsky, Soumith Chintala, and Léon Bottou · 2017
Later among the works it cites.
Why momentum really works
Gabriel Goh · 2017
Later among the works it cites.
Improved training of wasserstein gans
Ishaan Gulrajani, Faruk Ahmed, Martin Arjovsky, Vincent Dumoulin, and Aaron C Courville · 2017
Later among the works it cites.
Accelerated gradient descent escapes saddle points faster than gradient descent
Chi Jin, Praneeth Netrapalli, and Michael I Jordan · 2017
Later among the works it cites.
Sharpness, restart and acceleration
Vincent Roulet and Alexandre d’Aspremont · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Qihang Lin and Lin Xiao · 2014
Cited alongside, same era.
Efficient first-order methods for linear programming and semidefinite programming
James Renegar · 2014
Cited alongside, same era.
A differential equation for modeling nesterov’s accelerated gradient method: Theory and insights
Weijie Su, Stephen Boyd, and Emmanuel Candes · 2014
Cited alongside, same era.
A simple way to initialize recurrent networks of rectified linear units
Quoc V Le, Navdeep Jaitly, and Geoffrey E Hinton · 2015
Cited alongside, same era.
Adaptive restart for accelerated gradient schemes
Brendan O’donoghue and Emmanuel Candes · 2015
Cited alongside, same era.
Computational complexity versus statistical performance on sparse recovery problems
Vincent Roulet, Nicolas Boumal, and Alexandre d’Aspremont · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, et al · 2015
Cited alongside, same era.
Pytorch classification
Wei Yang · 2017
Later among the works it cites.
Robust accelerated gradient methods for smooth strongly convex functions
Necdet Serhat Aybat, Alireza Fallah, Mert Gurbuzbalaban, and Asuman Ozdaglar · 2018
Later among the works it cites.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2018
Later among the works it cites.
On acceleration with noise-corrupted gradients
Michael B Cohen, Jelena Diakonikolas, and Lorenzo Orecchia · 2018
Later among the works it cites.
New computational guarantees for solving convex optimization problems with first order methods, via a function growth condition measure
Robert M Freund and Haihao Lu · 2018
Later among the works it cites.
Fixing weight decay regularization in adam
Ilya Loshchilov and Frank Hutter · 2018
Later among the works it cites.
Laplacian smoothing gradient descent
Stanley Osher, Bao Wang, Penghang Yin, Xiyang Luo, Farzin Barekat, Minh Pham, and Alex Lin · 2018
Later among the works it cites.
Understanding generalization through visualizations
W Ronny Huang, Zeyad Emam, Micah Goldblum, Liam Fowl, Justin K Terry, Furong Huang, and Tom Goldstein · 2019
Later among the works it cites.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Later among the works it cites.
On the convergence of adam and beyond
Sashank J Reddi, Satyen Kale, and Sanjiv Kumar · 2019
Later among the works it cites.
On the variance of the adaptive learning rate and beyond
Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han · 2020
Closest in time.