Fetching the paper…
Reading the bibliography…
Momentum based stochastic gradient methods such as heavy ball (HB) and Nesterov's accelerated gradient descent (NAG) method are widely used in practice for training deep networks and other supervised learning models, as they often provide significant improvements over stochastic gradient descent (SGD).
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
B. T. Polyak · 1964
Earlier work this paper cites.
The computation of eigenvalues and eigenvectors of very large sparse matrices
C. C. Paige · 1971
Earlier work this paper cites.
Channel identification for high speed digital communications
J. G. Proakis · 1974
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate o (1/k2)
Y. Nesterov · 1983
Earlier work this paper cites.
Introduction to Optimization
B. T. Polyak · 1987
Earlier work this paper cites.
Behavior of slightly perturbed lanczos and conjugate-gradient recurrences
A. Greenbaum · 1989
Earlier work this paper cites.
Analysis of the momentum lms algorithm
S. Roy and J. J. Shynk · 1990
Earlier work this paper cites.
Analysis of momentum adaptive filtering algorithms
R. Sharma, W. A. Sethares, and J. A. Bucklew · 1998
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course , volume 87 of Applied Optimization
Y. E. Nesterov · 2004
Earlier work this paper cites.
Reducing the dimensionality of data with neural networks
G. E. Hinton and R. R. Salakhutdinov · 2006
Earlier work this paper cites.
The tradeoffs of large scale learning
L. Bottou and O. Bousquet · 2007
Earlier work this paper cites.
Smooth optimization with approximate gradient
A. d’Aspremont · 2008
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky and G. Hinton · 2009
Earlier work this paper cites.
Deep learning via hessian-free optimization
J. Martens · 2010
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
J. C. Duchi, E. Hazan, and Y. Singer · 2011
Cited alongside, same era.
Optimal stochastic approximation algorithms for strongly convex stochastic composite optimization i: A generic algorithmic framework
S. Ghadimi and G. Lan · 2012
Cited alongside, same era.
Making gradient descent optimal for strongly convex stochastic optimization
A. Rakhlin, O. Shamir, and K. Sridharan · 2012
Cited alongside, same era.
A stochastic gradient method with an exponential convergence rate for strongly-convex optimization with finite training sets
N. L. Roux, M. Schmidt, and F. R. Bach · 2012
Cited alongside, same era.
Stochastic dual coordinate ascent methods for regularized loss minimization
A universal catalyst for first-order optimization
H. Lin, J. Mairal, and Z. Harchaoui · 2015
Later among the works it cites.
Optimizing neural networks with kronecker-factored approximate curvature
J. Martens and R. Grosse · 2015
Later among the works it cites.
Path-sgd: Path-normalized optimization in deep neural networks
B. Neyshabur, R. Salakhutdinov, and N. Srebro · 2015
Later among the works it cites.
Katyusha: The first direct acceleration of stochastic gradient methods
Z. Allen-Zhu · 2016
Later among the works it cites.
A simple practical accelerated method for finite sums
A. Defazio · 2016
Later among the works it cites.
Harder, better, faster, stronger convergence rates for least-squares regression
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
S. Shalev-Shwartz and T. Zhang · 2012
Cited alongside, same era.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Cited alongside, same era.
Optimal stochastic approximation algorithms for strongly convex stochastic composite optimization, ii: shrinking procedures and optimal algorithms
S. Ghadimi and G. Lan · 2013
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
R. Johnson and T. Zhang · 2013
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Cited alongside, same era.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
A. Defazio, F. R. Bach, and S. Lacoste-Julien · 2014
Cited alongside, same era.
First-order methods of smooth convex optimization with inexact oracle
O. Devolder, F. Glineur, and Y. E. Nesterov · 2014
Cited alongside, same era.
A. Dieuleveut, N. Flammarion, and F. R. Bach · 2016
Later among the works it cites.
Parallelizing stochastic approximation through mini-batching and tail-averaging
P. Jain, S. M. Kakade, R. Kidambi, P. Netrapalli, and A. Sidford · 2016
Later among the works it cites.
Entropy-sgd: Biasing gradient descent into wide valleys
P. Chaudhari, A. Choromanska, S. Soatto, Y. LeCun, C. Baldassi, C. Borgs, J. Chayes, L. Sagun, and R. Zecchina · 2017
Later among the works it cites.
Accelerating stochastic gradient descent
P. Jain, S. M. Kakade, R. Kidambi, P. Netrapalli, and A. Sidford · 2017
Later among the works it cites.
Linearly convergent stochastic heavy ball method for minimizing generalization error
N. Loizou and P. Richtárik · 2017
Later among the works it cites.
Preresnet-44 for cifar-10
preresnet · 2017
Later among the works it cites.
A generic approach for escaping saddle points
S. Reddi, M. Zaheer, S. Sra, B. Poczos, F. Bach, R. Salakhutdinov, and A. Smola · 2017
Later among the works it cites.
Yellowfin and the art of momentum tuning
J. Zhang, I. Mitliagkas, and C. Ré · 2017
Later among the works it cites.