Fetching the paper…
Reading the bibliography…
Momentum is a simple and widely used trick which allows gradient-based optimizers to pick up speed along low curvature directions.
Some methods of speeding up the convergence of iteration methods
Boris T Polyak · 1964
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate o (1/k2)
Yurii Nesterov · 1983
Earlier work this paper cites.
Backpropagation through time: what it does and how to do it
Paul J Werbos · 1990
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Boris T Polyak and Anatoli B Juditsky · 1992
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Mitchell P Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini · 1993
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
Shun-Ichi Amari · 1998
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Martin Zinkevich · 2003
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Deep learning via hessian-free optimization
James Martens · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Classical mechanics
Herbert Goldstein · 2011
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Cited alongside, same era.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Cited alongside, same era.
Introductory lectures on convex optimization: A basic course , volume 87
Yurii Nesterov · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba · 2014
Cited alongside, same era.
Analysis and design of optimization algorithms via integral quadratic constraints
Laurent Lessard, Benjamin Recht, and Andrew Packard · 2016
Later among the works it cites.
A multi-step inertial forward-backward splitting method for non-convex optimization
Jingwei Liang, Jalal Fadili, and Gabriel Peyré · 2016
Later among the works it cites.
A variational perspective on accelerated methods in optimization
Andre Wibisono, Ashia C Wilson, and Michael I Jordan · 2016
Later among the works it cites.
A lyapunov analysis of momentum methods in optimization
Ashia C Wilson, Benjamin Recht, and Michael I Jordan · 2016
Later among the works it cites.
Why momentum really works
Gabriel Goh · 2017
Later among the works it cites.
The reversible residual network: Backpropagation without storing activations
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
James Martens · 2014
Cited alongside, same era.
A differential equation for modeling nesterov’s accelerated gradient method: Theory and insights
Weijie Su, Stephen Boyd, and Emmanuel Candes · 2014
Cited alongside, same era.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
Sergey Ioffe and Christian Szegedy · 2015
Cited alongside, same era.
Optimizing neural networks with kronecker-factored approximate curvature
James Martens and Roger Grosse · 2015
Cited alongside, same era.
Adaptive restart for accelerated gradient schemes
Brendan O’Donoghue and Emmanuel Candes · 2015
Cited alongside, same era.
On accelerated methods in optimization
Andre Wibisono and Ashia C Wilson · 2015
Cited alongside, same era.
Training deep nets with sublinear memory cost
Tianqi Chen, Bing Xu, Chiyuan Zhang, and Carlos Guestrin · 2016
Cited alongside, same era.
Aidan N Gomez, Mengye Ren, Raquel Urtasun, and Roger B Grosse · 2017
Later among the works it cites.
Regularizing and optimizing lstm language models
Stephen Merity, Nitish Shirish Keskar, and Richard Socher · 2017
Later among the works it cites.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer · 2017
Later among the works it cites.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Later among the works it cites.
On the insufficiency of existing momentum schemes for stochastic optimization
Rahul Kidambi, Praneeth Netrapalli, Prateek Jain, and Sham M Kakade · 2018
Closest in time.
On the convergence of adam and beyond
Sashank J. Reddi, Satyen Kale, and Sanjiv Kumar · 2018
Closest in time.
Adine: an adaptive momentum method for stochastic gradient descent
Vishwak Srinivasan, Adepu Ravi Sankar, and Vineeth N Balasubramanian · 2018
Closest in time.