Fetching the paper…
Reading the bibliography…
This paper studies a class of adaptive gradient based momentum algorithms that update the search directions and learning rates simultaneously using past gradients.
Some methods of speeding up the convergence of iteration methods
P. T. Polyak · 1964
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
A. Nemirovskii, D. B. Yudin, and E. R. Dawson · 1983
Earlier work this paper cites.
A method for unconstrained convex minimization problem with the rate of convergence o (1/kˆ 2)
Y. Nesterov · 1983
Earlier work this paper cites.
On the convergence of adagrad with momentum for training deep neural networks
F. Zou and L. Shen · 1983
Earlier work this paper cites.
Improving the convergence of back-propagation learning with second order methods
S. Becker, Y. Le Cun, et al · 1988
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
M. Zinkevich · 2003
Earlier work this paper cites.
Convex optimization
S. Boyd and L. Vandenberghe · 2004
Earlier work this paper cites.
On the complexity of steepest descent, newton’s and regularized newton’s methods for nonconvex unconstrained optimization problems
C. Cartis, N. I. Gould, and P. L. Toint · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Earlier work this paper cites.
Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude
T. Tieleman and G. Hinton · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
M. D. Zeiler · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
S. Ghadimi and G. Lan · 2013
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
R. Johnson and T. Zhang · 2013
Cited alongside, same era.
Introductory lectures on convex optimization: A basic course , volume 87
Y. Nesterov · 2013
Cited alongside, same era.
Adam: A method for stochastic optimization
D. P. Kingma and J. Ba · 2014
Cited alongside, same era.
Equilibrated adaptive learning rates for non-convex optimization
Y. Dauphin, H. de Vries, and Y. Bengio · 2015
Cited alongside, same era.
Global convergence of the heavy-ball method for convex optimization
E. Ghadimi, H. R. Feyzmahdavian, and M. Johansson · 2015
Accelerated gradient descent escapes saddle points faster than gradient descent
C. Jin, P. Netrapalli, and M. I. Jordan · 2017
Later among the works it cites.
Non-convex finite-sum optimization via scsg methods
L. Lei, C. Ju, J. Chen, and M. I. Jordan · 2017
Later among the works it cites.
A. Basu, S. De, A. Mukherjee, and E. Ullah · 2018
Closest in time.
signsgd: compressed optimisation for non-convex problems
J. Bernstein, Y. Wang, K. Azizzadenesheli, and A. Anandkumar · 2018
Closest in time.
Optimization methods for large-scale machine learning
L. Bottou, F. E. Curtis, and J. Nocedal · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
ipiasco: Inertial proximal algorithm for strongly convex optimization
P. Ochs, T. Brox, and T. Pock · 2015
Cited alongside, same era.
Incorporating nesterov momentum into adam
T. Dozat · 2016
Cited alongside, same era.
Accelerated gradient methods for nonconvex nonlinear and stochastic programming
S. Ghadimi and G. Lan · 2016
Cited alongside, same era.
Stochastic variance reduction for nonconvex optimization
S. J. Reddi, A. Hefny, S. Sra, B. Poczos, and A. Smola · 2016
Cited alongside, same era.
Unified convergence analysis of stochastic momentum methods for convex and non-convex optimization
T. Yang, Q. Lin, and Z. Li · 2016
Cited alongside, same era.
Closing the generalization gap of adaptive gradient methods in training deep neural networks
J. Chen and Q. Gu · 2018
Closest in time.
Nostalgic adam: Weighing more of the past gradients when designing the adaptive learning rate
H. Huang, C. Wang, and B. Dong · 2018
Closest in time.
On the convergence of stochastic gradient descent with adaptive stepsizes
X. Li and F. Orabona · 2018
Closest in time.
On the convergence of adam and beyond
S. J. Reddi, S. Kale, and S. Kumar · 2018
Closest in time.
Adagrad stepsizes: Sharp convergence over nonconvex landscapes, from any initialization
R. Ward, X. Wu, and L. Bottou · 2018
Closest in time.
On the convergence of adaptive gradient methods for nonconvex optimization
D. Zhou, Y. Tang, Z. Yang, Y. Cao, and Q. Gu · 2018
Closest in time.