Fetching the paper…
Reading the bibliography…
Simple stochastic momentum methods are widely used in machine learning optimization, but their good practical performance is at odds with an absence of theoretical guarantees of acceleration in the literature.
Angenäherte auflösung von systemen linearer gleichungen
Kaczmarz, S. M. (1937) · 1937
Earlier work this paper cites.
Methods of conjugate gradients for solving linear systems
Hestenes, M. R. and Stiefel, E. (1952) · 1952
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Polyak, B. (1964) · 1964
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate o ( 1 / k 2 ) o(1/k^{2})
Nesterov, Y. (1983) · 1983
Earlier work this paper cites.
On the real convergence rate of the conjugate gradient method
Strakos, Z. (1991) · 1991
Earlier work this paper cites.
Open questions in the convergence analysis of the Lanczos process for the real symmetric eigenvalue problem
Strakos, Z. and Greenbaum, A. (1992) · 1992
Earlier work this paper cites.
A randomized Kaczmarz algorithm with exponential convergence
Strohmer, T. and Vershynin, R. (2008) · 2008
Earlier work this paper cites.
Adaptive bound optimization for online convex optimization
McMahan, B. and Streeter, M. (2010) · 2010
Earlier work this paper cites.
Randomized Kaczmarz solver for noisy linear systems
Needell, D. (2010) · 2010
Earlier work this paper cites.
Cs726-lyapunov analysis and the heavy ball method
Recht, B. (2010) · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
Duchi, J., Hazan, E., and Singer, Y. (2011) · 2011
Earlier work this paper cites.
The masked sample covariance estimator: an analysis using matrix concentration inequalities
Chen, R. Y., Gittens, A., and Tropp, J. A. (2012) · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Krizhevsky, A., Sutskever, I., and Hinton, G. E. (2012) · 2012
Earlier work this paper cites.
Krylov subspace methods: principles and analysis
Liesen, J. and Strakoš, Z. (2013) · 2013
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: a Basic Course
Nesterov, Y. (2013) · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Sutskever, I., Martens, J., Dahl, G., and Hinton, G. (2013) · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J. (2014) · 2014
Earlier work this paper cites.
Stochastic gradient descent, weighted sampling, and the randomized Kaczmarz algorithm
Needell, D., Srebro, N., and Ward, R. (2014) · 2014
Earlier work this paper cites.
Paved with good intentions: Analysis of a randomized block Kaczmarz method
Needell, D. and Tropp, J. A. (2014) · 2014
Cited alongside, same era.
A geometric alternative to nesterov’s accelerated gradient descent
Bubeck, S., Lee, Y. T., and Singh, M. (2015) · 2015
Cited alongside, same era.
From averaging to acceleration, there is only a step-size
Flammarion, N. and Bach, F. (2015) · 2015
Cited alongside, same era.
Un-regularizing: approximate proximal point and faster stochastic algorithms for empirical risk minimization
Frostig, R., Ge, R., Kakade, S., and Sidford, A. (2015) · 2015
Cited alongside, same era.
An accelerated randomized Kaczmarz algorithm
Liu, J. and Wright, S. J. (2015) · 2015
Cited alongside, same era.
An introduction to matrix concentration inequalities
On the insufficiency of existing momentum schemes for stochastic optimization
Kidambi, R., Netrapalli, P., Jain, P., and Kakade, S. (2018) · 2018
Later among the works it cites.
Catalyst acceleration for first-order convex optimization: from theory to practice
Lin, H., Mairal, J., and Harchaoui, Z. (2018) · 2018
Later among the works it cites.
The power of interpolation: Understanding the effectiveness of SGD in modern over-parametrized learning
Ma, S., Bassily, R., and Belkin, M. (2018) · 2018
Later among the works it cites.
A unified analysis of stochastic momentum methods for deep learning
Yan, Y., Yang, T., Li, Z., Lin, Q., and Yang, Y. (2018) · 2018
Later among the works it cites.
Accelerated linear convergence of stochastic momentum methods in Wasserstein distances
Can, B., Gurbuzbalaban, M., and Zhu, L. (2019) · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tropp, J. A. (2015) · 2015
Cited alongside, same era.
The ASTRA toolbox: A platform for advanced algorithm development in electron tomography
van Aarle, W., Palenstijn, W. J., Beenhouwer, J. D., Altantzis, T., Bals, S., Batenburg, K. J., and Sijbers, J. (2015) · 2015
Cited alongside, same era.
A simple practical accelerated method for finite sums
Defazio, A. (2016) · 2016
Cited alongside, same era.
Accelerated gradient methods for nonconvex nonlinear and stochastic programming
Ghadimi, S. and Lan, G. (2016) · 2016
Cited alongside, same era.
Batched stochastic gradient descent with weighted sampling
Needell, D. and Ward, R. (2016) · 2016
Cited alongside, same era.
Linearly convergent stochastic heavy ball method for minimizing generalization error
Loizou, N. and Richtárik, P. (2017) · 2017
Cited alongside, same era.
The fastest known globally convergent first-order method for minimizing strongly convex functions
Van Scoy, B., Freeman, R. A., and Lynch, K. M. (2017) · 2017
Cited alongside, same era.
Understanding the role of momentum in stochastic gradient methods
Gitman, I., Lang, H., Zhang, P., and Xiao, L. (2019) · 2019
Later among the works it cites.
Fast and faster convergence of sgd for over-parameterized models and an accelerated perceptron
Vaswani, S., Bach, F., and Schmidt, M. (2019) · 2019
Later among the works it cites.
Adagrad stepsizes: Sharp convergence over nonconvex landscapes
Ward, R., Wu, X., and Bottou, L. (2019) · 2019
Later among the works it cites.
Direct acceleration of saga using sampled negative momentum
Zhou, K., Ding, Q., Shang, F., Cheng, J., Li, D., and Luo, Z.-Q. (2019) · 2019
Later among the works it cites.
Robust accelerated gradient methods for smooth strongly convex functions
Aybat, N. S., Fallah, A., Gurbuzbalaban, M., and Ozdaglar, A. (2020) · 2020
Later among the works it cites.
A simple convergence proof of adam and adagrad
Défossez, A., Bottou, L., Bach, F., and Usunier, N. (2020) · 2020
Later among the works it cites.
An improved analysis of stochastic gradient descent with momentum
Liu, Y., Gao, Y., and Yin, W. (2020) · 2020
Later among the works it cites.
Momentum and stochastic momentum for stochastic gradient, Newton, proximal point and subspace descent methods
Loizou, N. and Richtárik, P. (2020) · 2020
Later among the works it cites.
Randomized Kaczmarz with averaging
Moorman, J. D., Tu, T. K., Molitor, D., and Needell, D. (2020) · 2020
Later among the works it cites.
Matrix concentration for products
Huang, D., Niles-Weed, J., Tropp, J. A., and Ward, R. (2021) · 2021
Later among the works it cites.
Almost sure convergence rates for stochastic gradient descent and stochastic heavy ball
Sebbouh, O., Gower, R. M., and Defazio, A. (2021) · 2021
Later among the works it cites.
Trajectory of mini-batch momentum: Batch size saturation and convergence in high dimensions
Lee, K., Cheng, A., Paquette, E., and Paquette, C. (2022) · 2022
Closest in time.