Fetching the paper…
Reading the bibliography…
Many popular learning-rate schedules for deep neural networks combine a decaying trend with local perturbations that attempt to escape saddle points and bad local minima.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
B. T. Polyak · 1964
Earlier work this paper cites.
Statistical learning theory
V. Vapnik · 1998
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
S. Ghadimi and G. Lan · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Earlier work this paper cites.
Beyond the regret minimization barrier: optimal algorithms for stochastic strongly-convex optimization
E. Hazan and S. Kale · 2014
Earlier work this paper cites.
Global convergence of the heavy-ball method for convex optimization
E. Ghadimi, H. R. Feyzmahdavian, and M. Johansson · 2015
Earlier work this paper cites.
Beyond convexity: stochastic quasi-convex optimization
E. Hazan, K. Y. Levy, and S. Shalev-Shwartz · 2015
Earlier work this paper cites.
A nonmonotone learning rate strategy for SGD training of deep neural networks
N. S. Keskar and G. Saon · 2015
Earlier work this paper cites.
Accelerated gradient methods for nonconvex nonlinear and stochastic programming
S. Ghadimi and G. Lan · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
S. Zagoruyko and N. Komodakis · 2016
Cited alongside, same era.
Exponential decay sine wave learning rate for fast deep neural network training
W. An, H. Wang, Y. Zhang, and Q. Dai · 2017
Cited alongside, same era.
SGDR: Stochastic gradient descent with warm restarts
I. Loshchilov and F. Hutter · 2017
Cited alongside, same era.
Cyclical learning rates for training neural networks
L. N. Smith · 2017
Cited alongside, same era.
Stochastic heavy ball
S. Gadat, F. Panloup, S. Saadane, et al · 2018
Cited alongside, same era.
Towards flatter loss surface via nonmonotonic learning rate scheduling
S. Seong, Y. Lee, Y. Kee, D. Han, and J. Kim · 2018
Cited alongside, same era.
Stagewise training accelerates convergence of testing error over SGD
Z. Yuan, Y. Yan, R. Jin, and T. Yang · 2019
Later among the works it cites.
A. Defazio · 2020
Later among the works it cites.
The complexity of finding stationary points with stochastic gradient descent
Y. Drori and O. Shamir · 2020
Later among the works it cites.
An improved analysis of stochastic gradient descent with momentum
Y. Liu, Y. Gao, and W. Yin · 2020
Later among the works it cites.
Convergence of a stochastic gradient method with momentum for non-smooth non-convex optimization
V. Mai and M. Johansson · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A unified analysis of stochastic momentum methods for deep learning
Y. Yan, T. Yang, Z. Li, Q. Lin, and Y. Yang · 2018
Cited alongside, same era.
Universal stagewise learning for non-convex problems with convergence on averaged solutions
Z. Chen, Z. Yuan, J. Yi, B. Zhou, E. Chen, and T. Yang · 2019
Cited alongside, same era.
Relative deviation learning bounds and generalization with unbounded loss functions
C. Cortes, S. Greenberg, and M. Mohri · 2019
Cited alongside, same era.
A stochastic trust region algorithm based on careful step normalization
F. E. Curtis, K. Scheinberg, and R. Shi · 2019
Cited alongside, same era.
The step decay schedule: A near optimal, geometrically decaying learning rate procedure for least squares
R. Ge, S. M. Kakade, R. Kidambi, and P. Netrapalli · 2019
Cited alongside, same era.
Understanding the role of momentum in stochastic gradient methods
I. Gitman, H. Lang, P. Zhang, and L. Xiao · 2019
Cited alongside, same era.
B. Shi, W. J. Su, and M. I. Jordan · 2020
Later among the works it cites.
Learning with non-convex truncated losses by sgd
Y. Xu, S. Zhu, S. Yang, C. Zhang, R. Jin, and T. Yang · 2020
Later among the works it cites.
https://www.cs.toronto.edu/~kriz/cifar.html
The CIFAR data set · 2021
Closest in time.
A second look at exponential and cosine step sizes: Simplicity, adaptivity, and performance
X. Li, Z. Zhuang, and F. Orabona · 2021
Closest in time.
On the hyperparameters in stochastic gradient descent with momentum
B. Shi · 2021
Closest in time.
On the convergence of stochastic gradient descent with bandwidth-based step size
X. Wang and Y.-x. Yuan · 2021
Closest in time.
On the convergence of step decay step-size for stochastic optimization
X. Wang, S. Magnússon, and M. Johansson · 2021
Closest in time.