Fetching the paper…
Reading the bibliography…
Adaptive methods such as Adam and RMSProp are widely used in deep learning but are not well understood.
Increased rates of convergence through learning rate adaptation
Robert A Jacobs · 1988
Earlier work this paper cites.
Cubic regularization of Newton method and its global performance
Yurii Nesterov and Boris T Polyak · 2006
Earlier work this paper cites.
Adaptive bound optimization for online convex optimization
H. Brendan McMahan and Matthew J. Streeter · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Freedman’s inequality for matrix martingales
Joel Tropp · 2011
Earlier work this paper cites.
Lecture 6.5-RMSProp: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course , volume 87
Yurii Nesterov · 2013
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Escaping from saddle points — online stochastic gradient for tensor decomposition
Rong Ge, Furong Huang, Chi Jin, and Yang Yuan · 2015
Earlier work this paper cites.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Earlier work this paper cites.
Gradient descent only converges to minimizers
Jason D. Lee, Max Simchowitz, Michael I. Jordan, and Benjamin Recht · 2016
Cited alongside, same era.
The power of normalization: Faster evasion of saddle points
Kfir Y Levy · 2016
Cited alongside, same era.
An overview of gradient descent optimization algorithms
Sebastian Ruder · 2016
Cited alongside, same era.
Finding approximate local minima faster than gradient descent
Naman Agarwal, Zeyuan Allen-Zhu, Brian Bullins, Elad Hazan, and Tengyu Ma · 2017
Cited alongside, same era.
Adaptive learning rate via covariance matrix based preconditioning for deep neural networks
Yasutoshi Ida, Yasuhiro Fujiwara, and Sotetsu Iwamura · 2017
Cited alongside, same era.
Accelerated methods for nonconvex optimization
Y. Carmon, J. Duchi, O. Hinder, and A. Sidford · 2018
Later among the works it cites.
On the convergence of a class of adam-type algorithms for non-convex optimization
Xiangyi Chen, Sijia Liu, Ruoyu Sun, and Mingyi Hong · 2018
Later among the works it cites.
Escaping saddles with stochastic gradients
Hadi Daneshmand, Jonas Kohler, Aurelien Lucchi, and Thomas Hofmann · 2018
Later among the works it cites.
Escaping saddle points in constrained optimization
Aryan Mokhtari, Asuman Ozdaglar, and Ali Jadbabaie · 2018
Later among the works it cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern · 2018
Later among the works it cites.
Adagrad stepsizes: Sharp convergence over nonconvex landscapes, from any initialization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chi Jin, Rong Ge, Praneeth Netrapalli, Sham M. Kakade, and Michael I. Jordan · 2017
Cited alongside, same era.
Variants of RMSProp and Adagrad with logarithmic regret bounds
Mahesh Chandra Mukkamala and Matthias Hein · 2017
Cited alongside, same era.
The case for full-matrix adaptive regularization
Naman Agarwal, Brian Bullins, Xinyi Chen, Elad Hazan, Karan Singh, Cyril Zhang, and Yi Zhang · 2018
Cited alongside, same era.
Neon2: Finding local minima via first-order oracles
Zeyuan Allen-Zhu and Yuanzhi Li · 2018
Cited alongside, same era.
Dissecting Adam: The sign, magnitude and variance of stochastic gradients
Lukas Balles and Philipp Hennig · 2018
Cited alongside, same era.
A generic approach for escaping saddle points
Sashank Reddi, Manzil Zaheer, Suvrit Sra, Barnabas Poczos, Francis Bach, Ruslan Salakhutdinov, and Alex Smola
Cited in the paper.
On the convergence of Adam and beyond
Sashank J. Reddi, Satyen Kale, and Sanjiv Kumar
Cited in the paper.
Rachel Ward, Xiaoxia Wu, and Leon Bottou · 2018
Later among the works it cites.
First-order stochastic algorithms for escaping from saddle points in almost linear time
Yi Xu, Jing Rong, and Tianbao Yang · 2018
Later among the works it cites.
Adaptive methods for nonconvex optimization
Manzil Zaheer, Sashank Reddi, Devendra Sachan, Satyen Kale, and Sanjiv Kumar · 2018
Later among the works it cites.
On the convergence of adaptive gradient methods for nonconvex optimization
Dongruo Zhou, Yiqi Tang, Ziyan Yang, Yuan Cao, and Quanquan Gu · 2018
Later among the works it cites.
A sufficient condition for convergences of Adam and RMSProp
Fangyu Zou, Li Shen, Zequn Jie, Weizhong Zhang, and Wei Liu · 2018
Later among the works it cites.