Fetching the paper…
Reading the bibliography…
Adaptive optimization methods such as AdaGrad, RMSprop and Adam have been proposed to achieve a rapid training process with an element-wise scaling term on learning rates.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Building a large annotated corpus of english: The penn treebank
Mitchell P. Marcus, Mary Ann Marcinkiewicz, and Beatrice Santorini · 1993
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber · 1997
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann Lecun, Leon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
On the generalization ability of on-line learning algorithms
Nicoì Cesa-Bianchi, Alex Conconi, and Claudio Gentile · 2002
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky and Geoffrey Hinton · 2009
Earlier work this paper cites.
Convolutional deep belief networks for scalable unsupervised learning of hierarchical representations
Honglak Lee, Roger Grosse, Rajesh Ranganath, and Andrew Y Ng · 2009
Earlier work this paper cites.
Adaptive bound optimization for online convex optimization
H Brendan Mcmahan and Matthew Streeter · 2010
Cited alongside, same era.
Yuxin Wu and Kaiming He · 2010
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Cited alongside, same era.
RMSprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Lei Ba · 2015
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Densely connected convolutional networks
Gao Huang, Zhuang Liu, Laurens van der Maaten, and Kilian Q Weinberger · 2017
Later among the works it cites.
Improving generalization performance by switching from Adam to SGD
Nitish Shirish Keskar and Richard Socher · 2017
Later among the works it cites.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Later among the works it cites.
On the convergence of a class of Adam-type algorithms for non-convex optimization
Xiangyi Chen, Sijia Liu, Ruoyu Sun, and Mingyi Hong · 2018
Later among the works it cites.
On the convergence of adam and beyond
Sashank J. Reddi, Stayen Kale, and Sanjiv Kumar · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Learning personalized end-to-end goal-oriented dialog
Liangchen Luo, Wenhao Huang, Qi Zeng, Zaiqing Nie, and Xu Sun · 2019
Closest in time.