Fetching the paper…
Reading the bibliography…
Several recently proposed stochastic optimization methods that have been successfully used in training deep networks such as RMSProp, Adam, Adadelta, Nadam are based on using gradient updates scaled by square roots of exponential moving averages of squared past gradients.
Adaptive and self-confident on-line learning algorithms
Peter Auer and Claudio Gentile · 2000
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Martin Zinkevich · 2003
Earlier work this paper cites.
On the generalization ability of on-line learning algorithms
Nicolò Cesa-Bianchi, Alex Conconi, and Claudio Gentile · 2004
Earlier work this paper cites.
Adaptive bound optimization for online convex optimization
H. Brendan McMahan and Matthew J. Streeter · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John C. Duchi, Elad Hazan, and Yoram Singer · 2011
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Cited alongside, same era.
RmsProp: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Cited alongside, same era.
ADADELTA: An Adaptive Learning Rate Method
Matthew D. Zeiler · 2012
Cited alongside, same era.
Dropout: A simple way to prevent neural networks from overfitting
Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov · 2014
Later among the works it cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Later among the works it cites.
Incorporating Nesterov Momentum into Adam
Timothy Dozat · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…