Fetching the paper…
Reading the bibliography…
We analyze the dynamics of large batch stochastic gradient descent with momentum (SGD+M) on the least squares problem when both the number of samples and dimensions are large.
Some methods of speeding up the convergence of iteration methods
B.T. Polyak · 1964
Earlier work this paper cites.
Distribution of eigenvalues for some sets of random matrices
V. Marchenko and L. Pastur · 1967
Earlier work this paper cites.
Applied probability and queues , volume 51 of Applications of Mathematics (New York)
S. Asmussen · 2003
Earlier work this paper cites.
Introductory lectures on convex optimization
Y. Nesterov · 2004
Earlier work this paper cites.
"mnist" handwritten digit database, 2010
Y. LeCun, C. Cortes, and C. Burges · 2010
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Earlier work this paper cites.
A note on the Hanson-Wright inequality for random vectors with dependencies
R. Adamczak · 2015
Earlier work this paper cites.
Concentration inequalities for sampling without replacement
R. Bardenet and Odalric-Ambrym M · 2015
Earlier work this paper cites.
From averaging to acceleration, there is only a step-size
N. Flammarion and F. Bach · 2015
Earlier work this paper cites.
Katyusha: The first direct acceleration of stochastic gradient methods
Z. Allen-Zhu · 2017
Cited alongside, same era.
Big Batch SGD: Automated Inference using Adaptive Batch Sizes
S. De, A. Yadav, D. Jacobs, and T. Goldstein · 2017
Cited alongside, same era.
Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour , 2017
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Cited alongside, same era.
Stochastic heavy ball
S. Gadat, F. Panloup, and S. Saadane · 2018
Cited alongside, same era.
Accelerating Stochastic Gradient Descent for Least Squares Regression
P. Jain, S. Kakade, R. Kidambi, P. Netrapalli, and A. Sidford · 2018
Cited alongside, same era.
On the insufficiency of existing momentum schemes for stochastic optimization
High-Dimensional Probability: An Introduction with Applications in Data Science
R. Vershynin · 2018
Later among the works it cites.
Which algorithmic choices matter at which batch sizes? insights from a noisy quadratic model
G. Zhang, L. Li, Z. Nado, J. Martens, S. Sachdeva, G. Dahl, C. Shallue, and R. Grosse · 2019
Later among the works it cites.
Accelerating SGD with momentum for over-parameterized learning
C. Liu and M. Belkin · 2020
Later among the works it cites.
Momentum and stochastic momentum for stochastic gradient, newton, proximal point and subspace descent methods
N. Loizou and P. Richtarik · 2020
Later among the works it cites.
The role of memory in stochastic optimization
A. Orvieto, J. Kohler, and A. Lucchi · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Kidambi, P. Netrapalli, P. Jain, and S. Kakade · 2018
Cited alongside, same era.
The power of interpolation: Understanding the effectiveness of sgd in modern over-parametrized learning
S. Ma, R. Bassily, and M. Belkin · 2018
Cited alongside, same era.
An Empirical Model of Large-Batch Training , 2018
S. McCandlish, J. Kaplan, D. Amodei, and OpenAI Dota Team · 2018
Cited alongside, same era.
Don’t Decay the Learning Rate, Increase the Batch Size
S.L. Smith, P.-J. Kindermans, C. Ying, and Q. V. Le · 2018
Cited alongside, same era.
C. Paquette and E. Paquette · 2021
Later among the works it cites.
Sgd in the large: Average-case analysis, asymptotics, and stepsize criticality
C. Paquette, K. Lee, F. Pedregosa, and E. Paquette · 2021
Later among the works it cites.
A hitchhiker’s guide to momentum, 2021
F. Pedregosa · 2021
Later among the works it cites.
Almost sure convergence rates for stochastic gradient descent and stochastic heavy ball
O. Sebbouh, R. Gower, and A. Defazio · 2021
Later among the works it cites.