Fetching the paper…
Reading the bibliography…
In this paper, we propose a novel accelerated stochastic gradient method with momentum, which momentum is the weighted average of previous gradients.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
B.T. Polyak · 1964
Earlier work this paper cites.
Optimal stochastic search and adaptive momentum
Todd K. Leen and Genevieve B. Orr · 1993
Earlier work this paper cites.
Dynamics and Algorithms for Stochastic Search
Genevieve Beth Orr · 1996
Cited alongside, same era.
Adagrad stepsizes: sharp convergence over nonconvex landscapes
Rachel Ward, Xiaoxia Wu, and Léon Bottou · 2006
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Cited alongside, same era.
Adam: A Method for Stochastic Optimization
Diederik P. Kingma and Jimmy Ba · 2014
Later among the works it cites.
Optimization methods for large-scale machine learning
Léon. Bottou, Frank E. Curtis, and Jorge. Nocedal · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…