Fetching the paper…
Reading the bibliography…
We consider gradient descent with `momentum', a widely used method for loss function minimization in machine learning.
Some methods of speeding up the convergence of iteration methods
Boris T Polyak · 1964
Earlier work this paper cites.
Mechanics
L.D. Landau and E.M. Lifshitz · 1982
Earlier work this paper cites.
A method for unconstrained convex minimization problem with the rate of convergence o ( 1 / k 2 ) o(1/k^{2})
Yurii Nesterov · 1983
Earlier work this paper cites.
Learning representations by back-propagating error
D Rumelhart, G Hinton, and R Williams · 1986
Earlier work this paper cites.
Scientific Computing and Differential Equations: An Introduction to Numerical Methods
G.H. Golub and J.M. Ortega · 1992
Earlier work this paper cites.
Efficient backprop
Y. LeCun, L. Bottou, G. Orr, and K. Muller · 1998
Earlier work this paper cites.
The mnist database of handwritten digits
Yann LeCun · 1998
Earlier work this paper cites.
On the momentum term in gradient descent learning algorithms
Ning Qian · 1999
Earlier work this paper cites.
On the convergence of adaptive gradient methods for nonconvex optimization
Dongruo Zhou, Yiqi Tang, Ziyan Yang, Yuan Cao, and Quanquan Gu · 2001
Earlier work this paper cites.
Introduction to Numerical Analysis
J. Stoer and R. Bulirsch · 2002
Earlier work this paper cites.
Classical Mechanics: Point Particles and Relativity
W. Greiner · 2003
Earlier work this paper cites.
Classical dynamics of particles and systems
S.T. Thornton and J.B. Marion · 2004
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
A differential equation for modeling nesterov’s accelerated gradient method: Theory and insights
Weijie Su, Stephen Boyd, and Emmanuel Candes · 2014
Cited alongside, same era.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Cited alongside, same era.
Neural Networks and Deep Learning
M.A. Nielsen · 2015
Cited alongside, same era.
On accelerated methods in optimization
Andre Wibisono and Ashia C Wilson · 2015
Cited alongside, same era.
Incorporating nesterov momentum into adam
Timothy Dozat · 2016
Aggregated momentum: Stability through passive damping
James Lucas, Shengyang Sun, Richard Zemel, and Roger Grosse · 2018
Later among the works it cites.
Quasi-hyperbolic momentum and adam for deep learning
Jerry Ma and Denis Yarats · 2018
Later among the works it cites.
Adine: An adaptive momentum method for stochastic gradient descent
Vishwak Srinivasan, Adepu Ravi Sankar, and Vineeth N Balasubramanian · 2018
Later among the works it cites.
Decaying momentum helps neural network training
John Chen and Anastasios Kyrillidis · 2019
Later among the works it cites.
On empirical comparisons of optimizers for deep learning
Dami Choi, Christopher J Shallue, Zachary Nado, Jaehoon Lee, Chris J Maddison, and George E Dahl · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Cited alongside, same era.
An overview of gradient descent optimization algorithms
Sebastian Ruder · 2016
Cited alongside, same era.
A variational perspective on accelerated methods in optimization
Andre Wibisono, Ashia C Wilson, and Michael I Jordan · 2016
Cited alongside, same era.
A lyapunov analysis of momentum methods in optimization
Ashia C Wilson, Benjamin Recht, and Michael I Jordan · 2016
Cited alongside, same era.
On the influence of momentum acceleration on online learning
Kun Yuan, Bicheng Ying, and Ali H. Sayed · 2016
Cited alongside, same era.
Accelerated gradient descent escapes saddle points faster than gradient descent
Chi Jin, Praneeth Netrapalli, and Michael I Jordan · 2017
Cited alongside, same era.
Later among the works it cites.
Understanding the role of momentum in stochastic gradient methods
Igor Gitman, Hunter Lang, Pengchuan Zhang, and Lin Xiao · 2019
Later among the works it cites.
Algorithms for optimization
Mykel J Kochenderfer and Tim A Wheeler · 2019
Later among the works it cites.
Nikola B Kovachki and Andrew M Stuart · 2019
Later among the works it cites.
Adaptive gradient methods with dynamic bound of learning rate
Liangchen Luo, Yuanhao Xiong, Yan Liu, and Xu Sun · 2019
Later among the works it cites.
A high-bias, low-variance introduction to machine learning for physicists
Pankaj Mehta, Marin Bukov, Ching-Hao Wang, Alexandre GR Day, Clint Richardson, Charles K Fisher, and David J Schwab · 2019
Later among the works it cites.
On the convergence of adam and beyond
Sashank J Reddi, Satyen Kale, and Sanjiv Kumar · 2019
Later among the works it cites.
A survey of optimization methods from a machine learning perspective
Shiliang Sun, Zehui Cao, Han Zhu, and Jing Zhao · 2019
Later among the works it cites.
Calibrating the learning rate for adaptive gradient methods to improve generalization performance
Qianqian Tong, Guannan Liang, and Jinbo Bi · 2019
Later among the works it cites.
Global convergence of adaptive gradient methods for an over-parameterized neural network
Xiaoxia Wu, Simon S Du, and Rachel Ward · 2019
Later among the works it cites.
Gradient descent based optimization algorithms for deep learning models training
Jiawei Zhang · 2019
Later among the works it cites.
Lookahead optimizer: k k steps forward, 1 1 step back
Michael Zhang, James Lucas, Jimmy Ba, and Geoffrey E Hinton · 2019
Later among the works it cites.