Fetching the paper…
Reading the bibliography…
Adjusting the learning rate schedule in stochastic gradient methods is an important unresolved problem which requires tuning in practice.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Introductory lectures on convex programming volume i: Basic course
Y. Nesterov · 1998
Earlier work this paper cites.
Nonlinear programming
B. Dimitri · 1999
Earlier work this paper cites.
Numerical Optimization
S. Wright and J. Nocedal · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
F. Bach and E. Moulines · 2011
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
M. Zeiler · 2012
Cited alongside, same era.
Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes
O. Shamir and T. Zhang · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
D. Kingma and J. Ba · 2014
Cited alongside, same era.
Stochastic gradient descent, weighted sampling, and the randomized kaczmarz algorithm
D. Needell, R. Ward, and N. Srebro · 2014
Cited alongside, same era.
Optimization methods for large-scale machine learning
L. Bottou, F. Curtis, and J. Nocedal · 2016
Later among the works it cites.
Weight normalization: A simple reparameterization to accelerate training of deep neural networks
T. Salimans and D. Kingma · 2016
Later among the works it cites.
D. Alexandre and B. Francis · 2017
Later among the works it cites.
Accurate, large minibatch SGD: training imagenet in 1 hour
P. Goyal, P. Dollár, R. B. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Later among the works it cites.
The marginal value of adaptive gradient methods in machine learning
A. Wilson, R. Roelofs, M. Stern, N. Srebro, and B. Recht · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Cited alongside, same era.
Stochastic optimization with importance sampling for regularized loss minimization
P. Zhao and T. Zhang · 2015
Cited alongside, same era.
Gradient-based learning applied to document recognition
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner
Cited in the paper.
Efficient backprop
Y. LeCun, L. Bottou, G. Orr, and K. Müller
Cited in the paper.
Neural networks for machine learning-lecture 6a-overview of mini-batch gradient descent
G. Hinton N. Srivastava and K. Swersky
Cited in the paper.
Later among the works it cites.
Adagrad stepsizes: Sharp convergence over nonconvex landscapes
R. Ward, X. Wu, and L. Bottou · 2019
Closest in time.