Fetching the paper…
Reading the bibliography…
The Adam algorithm has become extremely popular for large-scale machine learning.
The mnist database of handwritten digits
LeCun, Y · 1998
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Zinkevich, M · 2003
Earlier work this paper cites.
Convex optimization
Boyd, S. and Vandenberghe, L · 2004
Earlier work this paper cites.
Logarithmic regret algorithms for online convex optimization
Hazan, E., Agarwal, A., and Kale, S · 2007
Earlier work this paper cites.
On the generalization ability of online strongly convex programming algorithms
Kakade, S. M. and Tewari, A · 2009
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
Adaptive bound optimization for online convex optimization
McMahan, H. B. and Streeter, M · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Online learning and online convex optimization
Shalev-Shwartz, S. et al · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Cited alongside, same era.
Adadelta: an adaptive learning rate method
Zeiler, M. D · 2012
Cited alongside, same era.
Beyond the regret minimization barrier: optimal algorithms for stochastic strongly-convex optimization
Hazan, E. and Kale, S · 2014
Cited alongside, same era.
Draw: a recurrent neural network for image generation
Gregor, K., Danihelka, I., Graves, A., Rezende, D. J., and Wierstra, D · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2015
Cited alongside, same era.
Skip-thought vectors
Kiros, R., Zhu, Y., Salakhutdinov, R. R., Zemel, R., Urtasun, R., Torralba, A., and Fidler, S · 2015
Introduction to online convex optimization
Hazan, E. et al · 2016
Later among the works it cites.
Empirical investigation of optimization algorithms in neural machine translation
Bahar, P., Alkhouli, T., Peter, J.-T., Brix, C. J.-S., and Ney, H · 2017
Later among the works it cites.
Stronger baselines for trustable results in neural machine translation
Denkowski, M. and Neubig, G · 2017
Later among the works it cites.
Variants of rmsprop and adagrad with logarithmic regret bounds
Mukkamala, M. C. and Hein, M · 2017
Later among the works it cites.
Basu, A., De, S., Mukherjee, A., and Ullah, E · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Show, attend and tell: Neural image caption generation with visual attention
Xu, K., Ba, J., Kiros, R., Cho, K., Courville, A., Salakhudinov, R., Zemel, R., and Bengio, Y · 2015
Cited alongside, same era.
Incorporating nesterov momentum into adam
Dozat, T · 2016
Cited alongside, same era.
On the convergence of a class of adam-type algorithms for non-convex optimization
Chen, X., Liu, S., Sun, R., and Hong, M
Cited in the paper.
Sadagrad: Strongly adaptive stochastic gradient methods
Chen, Z., Xu, Y., Chen, E., and Yang, T
Cited in the paper.
On the convergence of adam and beyond
Reddi, S. J., Kale, S., and Kumar, S · 2018
Later among the works it cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Shazeer, N. and Stern, M · 2018
Later among the works it cites.
Gadam: Genetic-evolutionary adam for deep neural network optimization
Zhang, J., Cui, L., and Gouza, F. B · 2018
Later among the works it cites.