Fetching the paper…
Reading the bibliography…
We propose a new variant of AMSGrad, a popular adaptive gradient based optimization algorithm widely used for training deep neural networks.
Some methods of speeding up the convergence of iteration methods
B.T. Polyak · 1964
Earlier work this paper cites.
A polynomial extrapolation method for finding limits and antilimits of vector sequences
Stan Cabay and LW Jackson · 1976
Earlier work this paper cites.
Extrapolating to the limit of a vector sequence
RP Eddy · 1979
Earlier work this paper cites.
Learning to forget: Continual prediction with LSTM
Felix A. Gers, Jürgen Schmidhuber, and Fred A. Cummins · 2000
Earlier work this paper cites.
Introductory Lectures on Convex Optimization - A Basic Course
Yurii E. Nesterov · 2004
Earlier work this paper cites.
An empirical evaluation of deep architectures on problems with many factors of variation
Hugo Larochelle, Dumitru Erhan, Aaron C. Courville, James Bergstra, and Yoshua Bengio · 2007
Earlier work this paper cites.
On accelerated proximal gradient methods for convex-concave optimization
Paul Tseng · 2008
Earlier work this paper cites.
Robust logitboost and adaptive base class (abc) logitboost
Ping Li · 2010
Earlier work this paper cites.
Adaptive bound optimization for online convex optimization
H. Brendan McMahan and Matthew J. Streeter · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John C. Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Anderson acceleration for fixed-point iterations
Homer F. Walker and Peng Ni · 2011
Earlier work this paper cites.
Online optimization with gradual variations
Chao-Kai Chiang, Tianbao Yang, Chia-Jung Lee, Mehrdad Mahdavi, Chi-Jen Lu, Rong Jin, and Shenghuo Zhu · 2012
Earlier work this paper cites.
Rmsprop: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Adadelta: an adaptive learning rate method
Matthew D Zeiler · 2012
Earlier work this paper cites.
Extrapolation methods: theory and practice
Claude Brezinski and M Redivo Zaglia · 2013
Earlier work this paper cites.
Stochastic first- and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Cited alongside, same era.
Speech recognition with deep recurrent neural networks
Alex Graves, Abdel-rahman Mohamed, and Geoffrey E. Hinton · 2013
Cited alongside, same era.
Playing atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves, Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Cited alongside, same era.
Optimization, learning, and games with predictable sequences
Alexander Rakhlin and Karthik Sridharan · 2013
Cited alongside, same era.
Generative adversarial nets
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron C. Courville, and Yoshua Bengio · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
A unified analysis of stochastic momentum methods for deep learning
Yan Yan, Tianbao Yang, Zhe Li, Qihang Lin, and Yi Yang · 2018
Later among the works it cites.
Adaptive methods for nonconvex optimization
Manzil Zaheer, Sashank J. Reddi, Devendra Singh Sachan, Satyen Kale, and Sanjiv Kumar · 2018
Later among the works it cites.
On the convergence of adaptive gradient methods for nonconvex optimization
Dongruo Zhou, Yiqi Tang, Ziyan Yang, Yuan Cao, and Quanquan Gu · 2018
Later among the works it cites.
On the convergence of adagrad with momentum for training deep neural networks
Fangyu Zou and Li Shen · 2018
Later among the works it cites.
Efficient full-matrix adaptive regularization
Naman Agarwal, Brian Bullins, Xinyi Chen, Elad Hazan, Karan Singh, Cyril Zhang, and Yi Zhang · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Diederik P. Kingma and Jimmy Ba · 2015
Cited alongside, same era.
Fast convergence of regularized learning in games
Vasilis Syrgkanis, Alekh Agarwal, Haipeng Luo, and Robert E. Schapire · 2015
Cited alongside, same era.
Incorporating nesterov momentum into adam
Timothy Dozat · 2016
Cited alongside, same era.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
End-to-end training of deep visuomotor policies
Sergey Levine, Chelsea Finn, Trevor Darrell, and Pieter Abbeel · 2016
Cited alongside, same era.
Accelerating online convex optimization via adaptive prediction
Mehryar Mohri and Scott Yang · 2016
Cited alongside, same era.
Faster rates for convex-concave games
Jacob D. Abernethy, Kevin A. Lai, Kfir Y. Levy, and Jun-Kun Wang · 2018
Cited alongside, same era.
On the convergence of A class of adam-type algorithms for non-convex optimization
Xiangyi Chen, Sijia Liu, Ruoyu Sun, and Mingyi Hong · 2019
Closest in time.
Universal stagewise learning for non-convex problems with convergence on averaged solutions
Zaiyi Chen, Zhuoning Yuan, Jinfeng Yi, Bowen Zhou, Enhong Chen, and Tianbao Yang · 2019
Closest in time.
Introduction to online convex optimization
Elad Hazan · 2019
Closest in time.
On the convergence of stochastic gradient descent with adaptive stepsizes
Xiaoyu Li and Francesco Orabona · 2019
Closest in time.
Optimistic mirror descent in saddle-point problems: Going the extra (gradient) mile
Panayotis Mertikopoulos, Bruno Lecouat, Houssam Zenati, Chuan-Sheng Foo, Vijay Chandrasekhar, and Georgios Piliouras · 2019
Closest in time.
Adagrad stepsizes: sharp convergence over nonconvex landscapes
Rachel Ward, Xiaoxia Wu, and Léon Bottou · 2019
Closest in time.
Toward communication efficient adaptive gradient method
Xiangyi Chen, Xiaoyun Li, and Ping Li · 2020
Closest in time.
On the convergence of adam and adagrad
Alexandre Défossez, Léon Bottou, Francis Bach, and Nicolas Usunier · 2020
Closest in time.
Regularized nonlinear acceleration
Damien Scieur, Alexandre d’Aspremont, and Francis Bach · 2020
Closest in time.
Towards better generalization of adaptive gradient methods
Yingxue Zhou, Belhal Karimi, Jinxing Yu, Zhiqiang Xu, and Ping Li · 2020
Closest in time.