Fetching the paper…
Reading the bibliography…
Stochastic gradient methods with momentum are widely used in applications and at the core of optimization subroutines in many popular machine learning libraries.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
B. T. Polyak · 1964
Earlier work this paper cites.
Stochastic analog of the conjugant-gradient method
A. M. Gupal and L. G. Bazhenov · 1972
Earlier work this paper cites.
The quasigradient method for the solving of the nonlinear programming problems
E. A. Nurminskii · 1973
Earlier work this paper cites.
Favorable classes of Lipschitz continuous functions in subgradient optimization
R. T. Rockafellar · 1982
Earlier work this paper cites.
A method for solving the convex programming problem with convergence rate O ( 1 / k 2 ) O(1/k^{2})
Y. Nesterov · 1983
Earlier work this paper cites.
Stochastic approximation method with gradient averaging for unconstrained problems
A. Ruszczynski and W. Syski · 1983
Earlier work this paper cites.
Strong and weak convexity of sets and functions
J.-P. Vial · 1983
Earlier work this paper cites.
Introduction to optimization
B. T. Polyak · 1987
Earlier work this paper cites.
A linearization method for nonsmooth stochastic programming problems
A. Ruszczyński · 1987
Earlier work this paper cites.
Convex analysis and minimization algorithms
J.-B. Hiriart-Urruty and C. Lemaréchal · 1993
Earlier work this paper cites.
Heavy-ball method in nonconvex optimization problems
S. Zavriev and F. Kostyuk · 1993
Earlier work this paper cites.
Stochastic generalized gradient method for nonconvex nonsmooth stochastic optimization
Y. M. Ermol’ev and V. Norkin · 1998
Earlier work this paper cites.
An incremental gradient (-projection) method with momentum term and adaptive stepsize rule
P. Tseng · 1998
Earlier work this paper cites.
Optimization of conditional value-at-risk
R. T. Rockafellar and S. Uryasev · 2000
Earlier work this paper cites.
Accelerated gradient methods for stochastic optimization and online learning
C. Hu, W. Pan, and J. T. Kwok · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro · 2009
Earlier work this paper cites.
Primal-dual subgradient methods for convex problems
Y. Nesterov · 2009
Cited alongside, same era.
Large-scale machine learning with stochastic gradient descent
L. Bottou · 2010
Cited alongside, same era.
Dual averaging methods for regularized stochastic learning and online optimization
L. Xiao · 2010
Cited alongside, same era.
Robust principal component analysis?
E. J. Candès, X. Li, Y. Ma, and J. Wright · 2011
Cited alongside, same era.
Angular synchronization by eigenvectors and semidefinite programming
A. Singer · 2011
Cited alongside, same era.
Imagenet classification with deep convolutional neural networks
A. Krizhevsky, I. Sutskever, and G. E. Hinton · 2012
Cited alongside, same era.
First-order methods in optimization
A. Beck · 2017
Later among the works it cites.
The proximal point method revisited
D. Drusvyatskiy · 2017
Later among the works it cites.
Why momentum really works
G. Goh · 2017
Later among the works it cites.
Densely connected convolutional networks
G. Huang, Z. Liu, L. Van Der Maaten, and K. Q. Weinberger · 2017
Later among the works it cites.
Solving (most) of a set of quadratic equalities: Composite optimization for robust phase retrieval
J. C. Duchi and F. Ruan · 2018
Later among the works it cites.
Stochastic methods for composite and weakly convex optimization problems
J. C. Duchi and F. Ruan · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
S. Ghadimi and G. Lan · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Cited alongside, same era.
iPiano: Inertial proximal algorithm for nonconvex optimization
P. Ochs, Y. Chen, T. Brox, and T. Pock · 2014
Cited alongside, same era.
Lectures on stochastic programming: modeling and theory
A. Shapiro, D. Dentcheva, and A. Ruszczyński · 2014
Cited alongside, same era.
Convex optimization: Algorithms and complexity
S. Bubeck · 2015
Cited alongside, same era.
Global convergence of the heavy-ball method for convex optimization
E. Ghadimi, H. R. Feyzmahdavian, and M. Johansson · 2015
Cited alongside, same era.
Stochastic heavy ball
S. Gadat, F. Panloup, and S. Saadane · 2018
Later among the works it cites.
Lectures on convex optimization
Y. Nesterov · 2018
Later among the works it cites.
A unified analysis of stochastic momentum methods for deep learning
Y. Yan, T. Yang, Z. Li, Q. Lin, and Y. Yang · 2018
Later among the works it cites.
Lower bounds for non-convex stochastic optimization
Y. Arjevani, Y. Carmon, J. C. Duchi, D. J. Foster, N. Srebro, and B. Woodworth · 2019
Later among the works it cites.
Stochastic (approximate) proximal point methods: Convergence, optimality, and adaptivity
H. Asi and J. C. Duchi · 2019
Later among the works it cites.
Low-rank matrix recovery with composite optimization: good conditioning and rapid convergence
V. Charisopoulos, Y. Chen, D. Davis, M. Díaz, L. Ding, and D. Drusvyatskiy · 2019
Later among the works it cites.
Stochastic model-based minimization of weakly convex functions
D. Davis and D. Drusvyatskiy · 2019
Later among the works it cites.
Proximally guided stochastic subgradient method for nonsmooth, nonconvex problems
D. Davis and B. Grimmer · 2019
Later among the works it cites.
Efficiency of minimizing compositions of convex functions and smooth maps
D. Drusvyatskiy and C. Paquette · 2019
Later among the works it cites.
Understanding the role of momentum in stochastic gradient methods
I. Gitman, H. Lang, P. Zhang, and L. Xiao · 2019
Later among the works it cites.
A single time-scale stochastic approximation method for nested stochastic optimization
S. Ghadimi, A. Ruszczyński, and M. Wang · 2020
Closest in time.