Fetching the paper…
Reading the bibliography…
Momentum has become a crucial component in deep learning optimizers, necessitating a comprehensive understanding of when and why it accelerates stochastic gradient descent (SGD).
Some methods of speeding up the convergence of iteration methods
B. T. Polyak · 1964
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei · 2009
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
S. Merity, C. Xiong, J. Bradbury, and R. Socher · 2016
Earlier work this paper cites.
Eigenvalues of the hessian in deep learning: Singularity and beyond
L. Sagun, L. Bottou, and Y. LeCun · 2016
Earlier work this paper cites.
Accurate, large minibatch SGD: Training Imagenet in 1 hour
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Earlier work this paper cites.
The effect of network width on the performance of large-batch training
L. Chen, H. Wang, J. Zhao, D. Papailiopoulos, and P. Koutris · 2018
Earlier work this paper cites.
On the insufficiency of existing momentum schemes for stochastic optimization
R. Kidambi, P. Netrapalli, P. Jain, and S. M. Kakade · 2018
Cited alongside, same era.
The power of interpolation: Understanding the effectiveness of SGD in modern over-parametrized learning
S. Ma, R. Bassily, and M. Belkin · 2018
Cited alongside, same era.
On the convergence of Adam and beyond
S. J. Reddi, S. Kale, and S. Kumar · 2019
Cited alongside, same era.
Measuring the effects of data parallelism on neural network training
C. J. Shallue, J. Lee, J. Antognini, J. Sohl-Dickstein, R. Frostig, and G. E. Dahl · 2019
Cited alongside, same era.
Momentum improves normalized SGD
A. Cutkosky and H. Mehta · 2020
Cited alongside, same era.
Gradient descent on neural networks typically occurs at the edge of stability
J. M. Cohen, S. Kaur, Y. Li, J. Z. Kolter, and A. Talwalkar · 2021
Later among the works it cites.
Understanding gradient descent on the edge of stability in deep learning
S. Arora, Z. Li, and A. Panigrahi · 2022
Later among the works it cites.
On the fast convergence of minibatch heavy ball momentum
R. Bollapragada, T. Chen, and R. Ward · 2022
Later among the works it cites.
Adaptive gradient methods at the edge of stability
J. M. Cohen, B. Ghorbani, S. Krishnan, N. Agarwal, S. Medapati, M. Badura, D. Suo, D. Cardoze, Z. Nado, G. E. Dahl, et al · 2022
Later among the works it cites.
Noise is not the main factor behind the gap between SGD and Adam on transformers, but sign descent might be
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Defazio · 2020
Cited alongside, same era.
An image is worth 16x16 words: Transformers for image recognition at scale
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, et al · 2020
Cited alongside, same era.
The two regimes of deep network training
G. Leclerc and A. Madry · 2020
Cited alongside, same era.
An improved analysis of stochastic gradient descent with momentum
Y. Liu, Y. Gao, and W. Yin · 2020
Cited alongside, same era.
Learning threshold neurons via the" edge of stability"
K. Ahn, S. Bubeck, S. Chewi, Y. T. Lee, F. Suarez, and Y. Zhang
Cited in the paper.
Understanding the unstable convergence of gradient descent
K. Ahn, J. Zhang, and S. Sra
Cited in the paper.
Don’t decay the learning rate, increase the batch size
S. L. Smith, P.-J. Kindermans, C. Ying, and Q. V. Le
Cited in the paper.
F. Kunstner, J. Chen, J. W. Lavington, and M. Schmidt · 2022
Later among the works it cites.
Analyzing sharpness along gd trajectory: Progressive sharpening and edge of stability
Z. Li, Z. Wang, and J. Li · 2022
Later among the works it cites.
The multiscale structure of neural network loss functions: The effect on optimization and origin
C. Ma, L. Wu, and L. Ying · 2022
Later among the works it cites.
Understanding edge-of-stability training dynamics with a minimalist example
X. Zhu, Z. Wang, X. Wang, M. Zhou, and R. Ge · 2022
Later among the works it cites.
Symbolic discovery of optimization algorithms
X. Chen, C. Liang, D. Huang, E. Real, K. Wang, Y. Liu, H. Pham, X. Dong, T. Luong, C.-J. Hsieh, et al · 2023
Closest in time.