Fetching the paper…
Reading the bibliography…
Adam and RMSProp are two of the most influential adaptive stochastic algorithms for training deep neural networks, which have been pointed out to be divergent even in the convex setting via a few simple counterexamples.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1985
Earlier work this paper cites.
Gradient-based learning applied to document recognition
Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner · 1998
Earlier work this paper cites.
Nonlinear programming
Dimitri P Bertsekas · 1999
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A Krizhevsky · 2009
Earlier work this paper cites.
Mnist handwritten digit database. 2010
Yann LeCun, Corinna Cortes, and Christopher JC Burges · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Neural networks for machine learning lecture 6a overview of mini-batch gradient descent
Geoffrey Hinton, Nitish Srivastava, and Kevin Swersky · 2012
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Incorporating Nesterov momentum into Adam
Timothy Dozat · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
Non-convex optimization for machine learning
Prateek Jain, Purushottam Kar, et al · 2017
Cited alongside, same era.
Variants of RMSProp and Adagrad with logarithmic regret bounds
Mahesh Chandra Mukkamala and Matthias Hein · 2017
Cited alongside, same era.
Dissecting Adam: The sign, magnitude and variance of stochastic gradients
Lukas Balles and Philipp Hennig · 2018
Cited alongside, same era.
Amitabh Basu, Soham De, Anirbit Mukherjee, and Enayat Ullah · 2018
Cited alongside, same era.
signSGD: Compressed optimisation for non-convex problems
Nostalgic Adam: Weighing more of the past gradients when designing the adaptive learning rate
Haiwen Huang, Chang Wang, and Bin Dong · 2018
Closest in time.
On the convergence of stochastic gradient descent with adaptive stepsizes
Xiaoyu Li and Francesco Orabona · 2018
Closest in time.
On the convergence of Adam and beyond
Sashank J. Reddi, Satyen Kale, and Sanjiv Kumar · 2018
Closest in time.
AdaGrad stepsizes: Sharp convergence over nonconvex landscapes, from any initialization
Rachel Ward, Xiaoxia Wu, and Leon Bottou · 2018
Closest in time.
Adaptive methods for nonconvex optimization
Manzil Zaheer, Sashank Reddi, Devendra Sachan, Satyen Kale, and Sanjiv Kumar · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jeremy Bernstein, Yu-Xiang Wang, Kamyar Azizzadenesheli, and Animashree Anandkumar · 2018
Cited alongside, same era.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2018
Cited alongside, same era.
On the convergence of a class of Adam-type algorithms for non-convex optimization
Xiangyi Chen, Sijia Liu, Ruoyu Sun, and Mingyi Hong · 2018
Cited alongside, same era.
Universal stagewise learning for non-convex problems with convergence on averaged solutions
Zaiyi Chen, Tianbao Yang, Jinfeng Yi, Bowen Zhou, and Enhong Chen · 2018
Cited alongside, same era.
Dongruo Zhou, Yiqi Tang, Ziyan Yang, Yuan Cao, and Quanquan Gu · 2018
Closest in time.
AdaShift: Decorrelation and convergence of adaptive learning rate methods
Zhiming Zhou, Qingru Zhang, Guansong Lu, Hongwei Wang, Weinan Zhang, and Yong Yu · 2018
Closest in time.
On the convergence of AdaGrad with momentum for training deep neural networks
Fangyu Zou and Li Shen · 2018
Closest in time.