Fetching the paper…
Reading the bibliography…
We study Nesterov's accelerated gradient method with constant step-size and momentum parameters in the stochastic approximation setting (unbiased gradients with bounded variance) and the finite-sum setting (where randomness is due to sampling mini-batches).
A note on the joint spectral radius
Rota, G.-C. and Strang, W · 1960
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Polyak, B. T · 1964
Earlier work this paper cites.
A method for solving a convex programming problem with convergence rate 𝒪 ( 1 / k 2 ) \mathcal{O}(1/k^{2})
Nesterov, Y · 1983
Earlier work this paper cites.
Randomly generated test problems for positive definite quadratic programming
Lenard, M. L. and Minkoff, M · 1984
Earlier work this paper cites.
Introduction to Optimization
Polyak, B. T · 1987
Earlier work this paper cites.
Stochastic dynamics of learning with momentum in neural networks
Wiegerinck, W., Komoda, A., and Heskes, T · 1994
Earlier work this paper cites.
Introductory lectures on convex optimization: a basic course
Nesterov, Y · 2004
Earlier work this paper cites.
Smooth optimization with approximate gradient
d’Aspremont, A · 2008
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E · 2011
Earlier work this paper cites.
Optimal stochastic approximation algorithms for strongly convex stochastic composite optimization I: A generic algorithmic framework
Ghadimi, S. and Lan, G · 2012
Earlier work this paper cites.
An optimal method for stochastic composite optimization
Lan, G · 2012
Earlier work this paper cites.
Optimal stochastic approximation algorithms for strongly convex stochastic composite optimization II: Shrinking procedures and optimal algorithms
Ghadimi, S. and Lan, G · 2013
Earlier work this paper cites.
Matrix Analysis
Horn, R. A. and Johnson, C. R · 2013
Earlier work this paper cites.
Fast convergence of stochastic gradient descent under a strong growth condition
Schmidt, M. and Le Roux, N · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Sutskever, I., Martens, J., Dahl, G., and Hinton, G · 2013
Cited alongside, same era.
Linear coupling: An ultimate unification of gradient and mirror descent
Allen-Zhu, Z. and Orecchia, L · 2014
Cited alongside, same era.
First-order methods of smooth convex optimization with inexact oracle
Devolder, O., Glineur, F., and Nesterov, Y · 2014
Cited alongside, same era.
A differential equation for modeling nesterov’s accelerated gradient method: Theory and insights
Su, W., Boyd, S., and Candès, E · 2014
Cited alongside, same era.
Convex optimization: Algorithms and complexity
Bubeck, S · 2015
Cited alongside, same era.
Adaptive restart for accelerated gradient schemes
O’Donoghue, B. and Candès, E · 2015
On acceleration with noise-corrupted gradients
Cohen, M. B., Diakonikolas, J., and Orecchia, L · 2018
Later among the works it cites.
Negative momentum for improved game dynamics
Gidel, G., Hemmat, R. A., Pezeshki, M., Lepriol, R., Huang, G., Lacoste-Julien, S., and Mitliagkas, I · 2018
Later among the works it cites.
On the insufficiency of existing momentum schemes for stochastic optimization
Kidambi, R., Netrapalli, P., Jain, P., and Kakade, S. M · 2018
Later among the works it cites.
The power of interpolation: Understanding the effectiveness of SGD in modern over-parameterized learning
Ma, S., Bassily, R., and Belkin, M · 2018
Later among the works it cites.
Robust accelerated gradient methods for smooth strongly convex functions
Aybat, N. S., Fallah, A., Gürbüzbalaban, M., and Ozdaglar, A · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Analysis and design of optimization algorithms via integral quadratic constraints
Lessard, L., Recht, B., and Packard, A · 2016
Cited alongside, same era.
Asynchrony begets momentum, with an application to deep learning
Mitliagkas, I., Zhang, C., Hadjis, S., and Ré, C · 2016
Cited alongside, same era.
Unified convergence analysis of stochastic momentum methods for convex and non-convex optimization
Yang, T., Lin, Q., and Li, Z · 2016
Cited alongside, same era.
Katyusha: The first direct acceleration of stochastic gradient methods
Allen-Zhu, Z · 2017
Cited alongside, same era.
Dissipativity theory for Nesterov’s accelerated method
Hu, B. and Lessard, L · 2017
Cited alongside, same era.
Loizou, N. and Richtárik, P · 2017
Cited alongside, same era.
Can, B., Gurbuzbalaban, M., and Zhu, L · 2019
Later among the works it cites.
On the curved geometry of accelerated optimization
Defazio, A · 2019
Later among the works it cites.
Understanding the role of momentum in stochastic gradient methods
Gitman, I., Lang, H., Zhang, P., and Xiao, L · 2019
Later among the works it cites.
SGD: General analysis and improved rates
Gower, R., Loizou, N., Qian, X., Sailanbayev, A., Shulgin, E., and Richtarik, P · 2019
Later among the works it cites.
A Lyapunov analysis for accelerated gradient methods: From deterministic to stochastic case
Laborde, M. and Oberman, A · 2019
Later among the works it cites.
Quasi-hyperbolic momentum and adam for deep learning
Ma, J. and Yarats, D · 2019
Later among the works it cites.
Fast and faster convergence of SGD for over-parameterized models (and an accelerated perceptron)
Vaswani, S., Bach, F., and Schmidt, M · 2019
Later among the works it cites.
Accelerating SGD with momentum for over-parameterized learning
Liu, C. and Belkin, M · 2020
Closest in time.