Fetching the paper…
Reading the bibliography…
While stochastic gradient descent (SGD) is still the \emph{de facto} algorithm in deep learning, adaptive methods like Clipped SGD/Adam have been observed to outperform SGD across important tasks, such as attention models.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Minimization of functions having Lipschitz continuous first partial derivatives
L. Armijo · 1966
Earlier work this paper cites.
Introduction to optimization. optimization software
B. T. Polyak · 1987
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
A. Rakhlin, O. Shamir, and K. Sridharan · 2011
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
Bandits with heavy tail
S. Bubeck, N. Cesa-Bianchi, and G. Lugosi · 2013
Earlier work this paper cites.
ADAM: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Earlier work this paper cites.
Beyond convexity: Stochastic quasi-convex optimization
E. Hazan, K. Levy, and S. Shalev-Shwartz · 2015
Earlier work this paper cites.
The power of normalization: Faster evasion of saddle points
K. Y. Levy · 2016
Earlier work this paper cites.
Lower bounds for finding stationary points i
Y. Carmon, J. C. Duchi, O. Hinder, and A. Sidford · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
The marginal value of adaptive gradient methods in machine learning
A. C. Wilson, R. Roelofs, M. Stern, N. Srebro, and B. Recht · 2017
Cited alongside, same era.
The case for full-matrix adaptive regularization
N. Agarwal, B. Bullins, X. Chen, E. Hazan, K. Singh, C. Zhang, and Y. Zhang · 2018
Cited alongside, same era.
On the convergence of a class of adam-type algorithms for non-convex optimization
X. Chen, S. Liu, R. Sun, and M. Hong · 2018
Cited alongside, same era.
Bert: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2018
Cited alongside, same era.
Nostalgic adam: Weighting more of the past gradients when designing the adaptive learning rate
On the adequacy of untuned warmup for adaptive optimization
J. Ma and D. Yarats · 2019
Closest in time.
First exit time analysis of stochastic gradient descent under heavy-tailed gradient noise
T. H. Nguyen, U. Şimşekli, M. Gürbüzbalaban, and G. Richard · 2019
Closest in time.
Non-gaussianity of stochastic gradient noise
A. Panigrahi, R. Somani, N. Goyal, and P. Netrapalli · 2019
Closest in time.
On the convergence of ADAM and beyond
S. J. Reddi, S. Kale, and S. Kumar · 2019
Closest in time.
A tail-index analysis of stochastic gradient noise in deep neural networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
H. Huang, C. Wang, and B. Dong · 2018
Cited alongside, same era.
On the convergence of stochastic gradient descent with adaptive stepsizes
X. Li and F. Orabona · 2018
Cited alongside, same era.
Adagrad stepsizes: Sharp convergence over nonconvex landscapes, from any initialization
R. Ward, X. Wu, and L. Bottou · 2018
Cited alongside, same era.
On the convergence of weighted adagrad with momentum for training deep neural networks
F. Zou and L. Shen · 2018
Cited alongside, same era.
Lower bounds for non-convex stochastic optimization
Y. Arjevani, Y. Carmon, J. C. Duchi, D. J. Foster, N. Srebro, and B. Woodworth · 2019
Cited alongside, same era.
Transformer-xl: Attentive language models beyond a fixed-length context
Z. Dai, Z. Yang, Y. Yang, J. Carbonell, Q. V. Le, and R. Salakhutdinov · 2019
Cited alongside, same era.
On the variance of the adaptive learning rate and beyond
L. Liu, H. Jiang, P. He, W. Chen, X. Liu, J. Gao, and J. Han · 2019
Cited alongside, same era.
On the convergence of adaptive gradient methods for nonconvex optimization
D. Zhou, Y. Tang, Z. Yang, Y. Cao, and Q. Gu
Cited in the paper.
U. Simsekli, L. Sagun, and M. Gurbuzbalaban · 2019
Closest in time.
Escaping saddle points with adaptive gradient methods
M. Staib, S. J. Reddi, S. Kale, S. Kumar, and S. Sra · 2019
Closest in time.
A sufficient condition for convergences of adam and rmsprop
F. Zou, L. Shen, Z. Jie, W. Zhang, and W. Liu · 2019
Closest in time.
Momentum improves normalized sgd
A. Cutkosky and H. Mehta · 2020
Closest in time.
Stochastic optimization with heavy-tailed noise via accelerated gradient clipping
E. Gorbunov, M. Danilova, and A. Gasnikov · 2020
Closest in time.
U. Şimşekli, L. Zhu, Y. W. Teh, and M. Gürbüzbalaban · 2020
Closest in time.
Why gradient clipping accelerates training: A theoretical justification for adaptivity
J. Zhang, T. He, S. Sra, and A. Jadbabaie · 2020
Closest in time.