Fetching the paper…
Reading the bibliography…
Adaptive gradient methods have become recently very popular, in particular as they have been shown to be useful in the training of deep neural networks.
Online convex programming and generalized infinitesimal gradient ascent
Zinkevich, M · 2003
Earlier work this paper cites.
Logarithmic regret algorithms for online convex optimization
Hazan, E., Agarwal, A., and Kale, S · 2007
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Bottou, L · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, John, Hazan, Elad, and Singer, Yoram · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y · 2011
Earlier work this paper cites.
Lecture 6d - a separate, adaptive learning rate for each connection
Hinton, G., Srivastava, N., and Swersky, K · 2012
Earlier work this paper cites.
ADADELTA: An adaptive learning rate method
Zeiler, M. D · 2012
Cited alongside, same era.
Beyond the regret minimization barrier: optimal algorithms for stochastic strongly-convex optimization
Hazan, E. and Kale, S · 2014
Cited alongside, same era.
A survey of algorithms and analysis for adaptive online learning
McMahan, H Brendan · 2014
Cited alongside, same era.
Unit tests for stochastic optimization
Schaul, T., Antonoglou, I., and Silver, D · 2014
Cited alongside, same era.
Equilibrated adaptive learning rates for non-convex optimization
Dauphin, Y., de Vries, H., and Bengio, Y · 2015
Cited alongside, same era.
Adam: a method for stochastic optimization
Deep learning in neural networks: An overview
Schmidhuber, J · 2015
Later among the works it cites.
Learning step size controllers for robust neural network training
Daniel, C., Taylor, J., and Nowozin, S · 2016
Later among the works it cites.
Introduction to online convex optimization
Hazan, E · 2016
Later among the works it cites.
Deep visual-semantic alignments for generating image descriptions
Karpathy, A. and Fei-Fei, L · 2016
Later among the works it cites.
An overview of gradient descent optimization algorithms
Ruder, S · 2016
Later among the works it cites.
A unified approach to adaptive regularization in online and stochastic optimization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Kingma, D. P. and Bai, J. L · 2015
Cited alongside, same era.
Deep residual learning for image recognition
He, Kaiming, Zhang, Xiangyu, Ren, Shaoqing, and Sun, Jian
Cited in the paper.
Identity mappings in deep residual networks
He, Kaiming, Zhang, Xiangyu, Ren, Shaoqing, and Sun, Jian
Cited in the paper.
Inception-v4, inception-resnet and the impact of residual connections on learning
Szegedy, C., Ioffe, S., and Vanhoucke, V
Cited in the paper.
Rethinking the inception architecture for computer vision
Szegedy, C., Vanhoucke, V., Ioffe, S., Shlens, J., and Wojna, Z
Cited in the paper.
Gupta, Vineet, Koren, Tomer, and Singer, Yoram · 2017
Closest in time.