Fetching the paper…
Reading the bibliography…
The convergence of stochastic gradient descent is highly dependent on the step-size, especially on non-convex problems such as neural network training.
Stochastic algorithms with geometric step decay converge linearly on sharp functions
D. Davis, D. Drusvyatskiy, and V. Charisopoulos · 1907
Earlier work this paper cites.
From low probability to high confidence in stochastic convex optimization
D. Davis, D. Drusvyatskiy, L. Xiao, and J. Zhang · 1907
Earlier work this paper cites.
Simple and optimal high-probability bounds for strongly-convex stochastic gradient descent
N. J. Harvey, C. Liaw, and S. Randhawa · 1909
Earlier work this paper cites.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
On convergence rates of subgradient optimization methods
J.-L. Goffin · 1977
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate o (1/kˆ 2)
Y. E. Nesterov · 1983
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course , volume 87
Y. Nesterov · 2004
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro · 2009
Earlier work this paper cites.
On the complexity of steepest descent, newton’s and regularized newton’s methods for nonconvex unconstrained optimization problems
C. Cartis, N. I. Gould, and P. L. Toint · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
E. Moulines and F. R. Bach · 2011
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Earlier work this paper cites.
S. Lacoste-Julien, M. Schmidt, and F. Bach · 2012
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
A. Rakhlin, O. Shamir, and K. Sridharan · 2012
Cited alongside, same era.
Minimization methods for non-differentiable functions , volume 3
N. Z. Shor · 2012
Cited alongside, same era.
Lecture 6.5-rmsprop, coursera: Neural networks for machine learning
T. Tieleman and G. Hinton · 2012
Cited alongside, same era.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
S. Ghadimi and G. Lan · 2013
Cited alongside, same era.
Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes
O. Shamir and T. Zhang · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Densely connected convolutional networks
G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger · 2017
Later among the works it cites.
Automatic differentiation in PyTorch
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer · 2017
Later among the works it cites.
The step decay schedule: A near optimal, geometrically decaying learning rate procedure for least squares
R. Ge, S. M. Kakade, R. Kidambi, and P. Netrapalli · 2019
Later among the works it cites.
SGD: General analysis and improved rates
R. M. Gower, N. Loizou, X. Qian, A. Sailanbayev, E. Shulgin, and P. Richtárik · 2019
Later among the works it cites.
Using statistics to automate stochastic optimization
H. Lang, L. Xiao, and P. Zhang · 2019
Later among the works it cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Beyond the regret minimization barrier: optimal algorithms for stochastic strongly-convex optimization
E. Hazan and S. Kale · 2014
Cited alongside, same era.
Tensorflow: Large-scale machine learning on heterogeneous distributed systems
M. Abadi, A. Agarwal, P. Barham, E. Brevdo, Z. Chen, C. Citro, G. Corrado, A. Davis, J. Dean, M. Devin, et al · 2015
Cited alongside, same era.
ADAM: A method for stochastic optimization
D. P. Kingma and J. L. Ba · 2015
Cited alongside, same era.
On the number of iterations for dantzig–wolfe optimization and packing-covering approximation algorithms
P. Klein and N. E. Young · 2015
Cited alongside, same era.
Mini-batch stochastic approximation methods for nonconvex stochastic composite optimization
S. Ghadimi, G. Lan, and H. Zhang · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
Painless stochastic gradient: Interpolation, line-search, and convergence rates
S. Vaswani, A. Mishkin, I. Laradji, M. Schmidt, G. Gidel, and S. Lacoste-Julien · 2019
Later among the works it cites.
Stagewise training accelerates convergence of testing error over SGD
Z. Yuan, Y. Yan, R. Jin, and T. Yang · 2019
Later among the works it cites.
The complexity of finding stationary points with stochastic gradient descent
Y. Drori and O. Shamir · 2020
Later among the works it cites.
A second look at exponential and cosine step sizes: Simplicity, convergence, and performance
X. Li, Z. Zhuang, and F. Orabona · 2020
Later among the works it cites.
Stochastic Polyak step-size for SGD: A adaptive learning rate for fast convergence
N. Loizou, S. Vaswani, I. Laradji, and S. Lacoste-Julien · 2020
Later among the works it cites.
Statistical adaptive stochastic gradient methods
P. Zhang, H. Lang, Q. Liu, and L. Xiao · 2020
Later among the works it cites.