Fetching the paper…
Reading the bibliography…
Stochastic gradient descent (\textsc{Sgd}) methods are the most powerful optimization tools in training machine learning and deep learning models.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
B. T. Polyak · 1964
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
A. Nemirovski and D. B. Yudin · 1983
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate 𝒪 ( 1 / k 2 ) \mathcal{O}(1/k^{2})
Y. Nesterov · 1983
Earlier work this paper cites.
New stochastic approximation type procedures
B. T. Polyak · 1990
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
B. T. Polyak and A. B. Juditsky · 1992
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course
Y. Nesterov · 2003
Earlier work this paper cites.
Cubic regularization of newton method and its global performance
Y. Nesterov and B. T. Polyak · 2006
Earlier work this paper cites.
On accelerated proximal gradient methods for convex-concave optimization
P. Tseng · 2008
Earlier work this paper cites.
Convex optimization under inexact first-order information
G. Lan · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro · 2009
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
J. C. Duchi, E. Hazan, and Y. Singer · 2011
Cited alongside, same era.
An optimal method for stochastic composite optimization
G. Lan · 2012
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Cited alongside, same era.
Adadelta: an adaptive learning rate method
M. D. Zeiler · 2012
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Incorporating nesterov momentum into adam
T. Dozat · 2016
Later among the works it cites.
Accelerated gradient methods for nonconvex nonlinear and stochastic programming
S. Ghadimi and G. Lan · 2016
Later among the works it cites.
Gradient descent only converges to minimizers
J. D. Lee, M. Simchowitz, M. I. Jordan, and B. Recht · 2016
Later among the works it cites.
Finding approximate local minima faster than gradient descent
N. Agarwal, Z. Allen-Zhu, B. Bullins, E. Hazan, and T. Ma · 2017
Later among the works it cites.
Accelerating stochastic gradient descent
P. Jain, S. M. Kakade, R. Kidambi, P. Netrapalli, and A. Sidford · 2017
Later among the works it cites.
Improving generalization performance by switching from adam to sgd
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Escaping from saddle points-online stochastic gradient for tensor decomposition
R. Ge, F. Huang, C. Jin, and Y. Yuan · 2015
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Cited alongside, same era.
Accelerated methods for non-convex optimization
Y. Carmon, J. C. Duchi, O. Hinder, and A. Sidford · 2016
Cited alongside, same era.
N. S. Keskar and R. Socher · 2017
Later among the works it cites.
Stochastic cubic regularization for fast nonconvex optimization
N. Tripuraneni, M. Stern, C. Jin, J. Regier, and M. I. Jordan · 2017
Later among the works it cites.
The marginal value of adaptive gradient methods in machine learning
A. C. Wilson, R. Roelofs, M. Stern, N. Srebro, and B. Recht · 2017
Later among the works it cites.
On the insufficiency of existing momentum schemes for stochastic optimization
R. Kidambi, P. Netrapalli, P. Jain, and S. M. Kakade · 2018
Closest in time.
On the convergence of adam and beyond
S. J. Reddi, S. Kale, and S. Kumar · 2018
Closest in time.