Fetching the paper…
Reading the bibliography…
The ADAM optimizer is exceedingly popular in the deep learning community.
A stochastic approximation method
Robbins, H. and Monro, S · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Polyak, B. T · 1964
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate 𝒪 ( 1 / k 2 ) \mathcal{O}(1/k^{2})
Nesterov, Y · 1983
Earlier work this paper cites.
The subgroup algorithm for generating uniform random variables
Diaconis, P. and Shahshahani, M · 1987
Earlier work this paper cites.
Improving the convergence of back-propagation learning with second order methods
Becker, S. and LeCun, Y · 1988
Earlier work this paper cites.
A direct adaptive method for faster backpropagation learning: The RPROP algorithm
Riedmiller, M. and Braun, H · 1993
Earlier work this paper cites.
Long short-term memory
Hochreiter, S. and Schmidhuber, J · 1997
Earlier work this paper cites.
Natural gradient works efficiently in learning
Amari, S.-I · 1998
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Cited alongside, same era.
RMSPROP: Divide the gradient by a running average of its recent magnitude
Tieleman, T. and Hinton, G · 2012
Cited alongside, same era.
ADADELTA: An adaptive learning rate method
Zeiler, M. D · 2012
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, R. and Zhang, T · 2013
Cited alongside, same era.
No more pesky learning rates
Schaul, T., Zhang, S., and LeCun, Y · 2013
Cited alongside, same era.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
Defazio, A., Bach, F., and Lacoste-Julien, S · 2014
ADAM: A method for stochastic optimization
Kingma, D. and Ba, J · 2015
Later among the works it cites.
Probabilistic line searches for stochastic optimization
Mahsereci, M. and Hennig, P · 2015
Later among the works it cites.
Linear convergence of gradient and proximal-gradient methods under the Polyak-Lojasiewicz condition
Karimi, H., Nutini, J., and Schmidt, M · 2016
Later among the works it cites.
Wide residual networks
Zagoruyko, S. and Komodakis, N · 2016
Later among the works it cites.
Automizing stochastic optimization with gradient variance estimates
Balles, L., Mahsereci, M., and Hennig, P · 2017
Closest in time.
Entropy-SGD: Biasing gradient descent into wide valleys
Chaudhari, P., Choromanska, A., Soatto, S., and LeCun, Y · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
New insights and perspectives on the natural gradient method
Martens, J · 2014
Cited alongside, same era.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs
Seide, F., Fu, H., Droppo, J., Li, G., and Yu, D · 2014
Cited alongside, same era.
TensorFlow: Large-scale machine learning on heterogeneous systems, 2015
Abadi, M., Agarwal, A., Barham, P., Brevdo, E., Chen, Z., Citro, C., Corrado, G. S., Davis, A., Dean, J., Devin, M., Ghemawat, S., Goodfellow, I., Harp, A., Irving, G., Isard, M., Jia, Y., Jozefowicz, R., Kaiser, L., Kudlur, M., Levenberg, J., Mané, D., Monga, R., Moore, S., Murray, D., Olah, C., Schuster, M., Shlens, J., Steiner, B., Sutskever, I., Talwar, K., Tucker, P., Vanhoucke, V., Vasudevan, V., Viégas, F., Vinyals, O., Warden, P., Wattenberg, M., Wicke, M., Yu, Y., and Zheng, X · 2015
Cited alongside, same era.
Coupling adaptive batch sizes with learning rates
Balles, L., Romero, J., and Hennig, P
Cited in the paper.
The marginal value of adaptive gradient methods in machine learning
Wilson, A. C., Roelofs, R., Stern, M., Srebro, N., and Recht, B · 2017
Closest in time.
Fashion-MNIST: A novel image dataset for benchmarking machine learning algorithms
Xiao, H., Rasul, K., and Vollgraf, R · 2017
Closest in time.