Fetching the paper…
Reading the bibliography…
Although ADAM is a very popular algorithm for optimizing the weights of neural networks, it has been recently shown that it can diverge even in simple convex optimization examples.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Une propriété topologique des sous-ensembles analytiques réels
S. Łojasiewicz · 1963
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
B. T. Polyak · 1964
Earlier work this paper cites.
On gradients of functions definable in o-minimal structures
K. Kurdyka · 1998
Earlier work this paper cites.
Introductory lectures on convex optimization: a basic course
Y. Nesterov · 2004
Earlier work this paper cites.
Convergence of the iterates of descent methods for analytic cost functions
P-A. Absil, R. Mahony, and B. Andrews · 2005
Earlier work this paper cites.
On the convergence of the proximal algorithm for nonsmooth functions involving analytic features
H. Attouch and J. Bolte · 2009
Earlier work this paper cites.
Proximal alternating minimization and projection methods for nonconvex problems: An approach based on the kurdyka-łojasiewicz inequality
H. Attouch, J. Bolte, P. Redont, and A. Soubeyran · 2010
Earlier work this paper cites.
Characterizations of łojasiewicz inequalities: subgradient flows, talweg, convexity
J. Bolte, A. Daniilidis, O. Ley, and L. Mazet · 2010
Earlier work this paper cites.
Adaptive bound optimization for online convex optimization
H. B. McMahan and M. J. Streeter · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
S. Ghadimi and G. Lan · 2013
Earlier work this paper cites.
On the difficulty of training recurrent neural networks
R. Pascanu, T. Mikolov, and Y. Bengio · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
I. Sutskever, J. Martens, G. Dahl, and G. Hinton · 2013
Earlier work this paper cites.
Proximal alternating linearized minimization for nonconvex and nonsmooth problems
J. Bolte, S. Sabach, and M. Teboulle · 2014
Earlier work this paper cites.
ipiano: Inertial proximal algorithm for nonconvex optimization
P. Ochs, Y. Chen, T. Brox, and T. Pock · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2015
Earlier work this paper cites.
Incorporating nesterov momentum into adam
T. Dozat · 2016
Cited alongside, same era.
Linear convergence of gradient and proximal-gradient methods under the polyak-łojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Cited alongside, same era.
A multi-step inertial forward-backward splitting method for non-convex optimization
J. Liang, J. Fadili, and G. Peyré · 2016
Cited alongside, same era.
A unified approach to adaptive regularization in online and stochastic optimization
V. Gupta, T. Koren, and Y. Singer · 2017
Cited alongside, same era.
Convergence rates of inertial splitting schemes for nonconvex composite optimization
P. R. Johnstone and P. Moulin · 2017
Cited alongside, same era.
Convergence analysis of proximal gradient with momentum for nonconvex optimization
An inertial newton algorithm for deep learning
C. Castera, J. Bolte, C. Févotte, and E. Pauwels · 2019
Closest in time.
On the convergence of a class of adam-type algorithms for non-convex optimization
X. Chen, S. Liu, R. Sun, and M. Hong · 2019
Closest in time.
Stochastic subgradient method converges on tame functions
D. Davis, D. Drusvyatskiy, S. Kakade, and J.D. Lee · 2019
Closest in time.
Generalized momentum-based methods: A hamiltonian perspective
J. Diakonikolas and M. I. Jordan · 2019
Closest in time.
On the convergence of stochastic gradient descent with adaptive stepsizes
X. Li and F. Orabona · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Q. Li, Y. Zhou, Y. Liang, and P. K. Varshney · 2017
Cited alongside, same era.
First order methods beyond convexity and lipschitz gradient continuity with applications to quadratic inverse problems
J. Bolte, S. Sabach, M. Teboulle, and Y. Vaisbourd · 2018
Cited alongside, same era.
Optimization methods for large-scale machine learning
L. Bottou, F. Curtis, and J. Nocedal · 2018
Cited alongside, same era.
Closing the generalization gap of adaptive gradient methods in training deep neural networks
J. Chen, D. Zhou, Y. Tang, Z. Yang, and Q. Gu · 2018
Cited alongside, same era.
S. De, A. Mukherjee, and E. Ullah · 2018
Cited alongside, same era.
Calculus of the exponent of kurdyka–łojasiewicz inequality and its applications to linear convergence of first-order methods
G. Li and T. K. Pong · 2018
Cited alongside, same era.
Local convergence of the heavy-ball method and ipiano for non-convex optimization
P. Ochs · 2018
Cited alongside, same era.
L. Liu, H. Jiang, P. He, W. Chen, X. Liu, J. Gao, and J. Han · 2019
Closest in time.
Adaptive gradient methods with dynamic bound of learning rate
L. Luo, Y. Xiong, and Y. Liu · 2019
Closest in time.
Quasi-hyperbolic momentum and adam for deep learning
J. Ma and D. Yarats · 2019
Closest in time.
On the convergence of adabound and its connection to sgd
P. Savarese · 2019
Closest in time.
Escaping saddle points with adaptive gradient methods
M. Staib, S. Reddi, S. Kale, S. Kumar, and S. Sra · 2019
Closest in time.
Adagrad stepsizes: Sharp convergence over nonconvex landscapes
R. Ward, X. Wu, and L. Bottou · 2019
Closest in time.
General inertial proximal gradient method for a class of nonconvex nonsmooth optimization problems
Z. Wu and M. Li · 2019
Closest in time.
Linear convergence of adaptive stochastic gradient descent
Y. Xie, X. Wu, and R. Ward · 2019
Closest in time.
Global convergence of block coordinate descent in deep learning
J. Zeng, T. T. Lau, S. Lin, and Y. Yao · 2019
Closest in time.
Why gradient clipping accelerates training: A theoretical justification for adaptivity
J. Zhang, T. He, S. Sra, and A. Jadbabaie · 2019
Closest in time.
Adashift: Decorrelation and convergence of adaptive learning rate methods
Z. Zhou, Q. Zhang, G. Lu, H. Wang, W. Zhang, and Y. Yu · 2019
Closest in time.
A sufficient condition for convergences of adam and rmsprop
F. Zou, L. Shen, Z. Jie, W. Zhang, and W. Liu · 2019
Closest in time.