Fetching the paper…
Reading the bibliography…
Adam is a commonly used stochastic optimization algorithm in machine learning.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Advanced Probability Theory , volume 10
Janos Galambos · 1995
Earlier work this paper cites.
Gradient convergence in gradient methods with errors
Dimitri P Bertsekas and John N Tsitsiklis · 2000
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A Basic Course , volume 87
Yurii Nesterov · 2003
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Arkadi Nemirovski, Anatoli B. Juditsky, Guanghui Lan, and Alexander Shapiro · 2008
Earlier work this paper cites.
Information-theoretic lower bounds on the oracle complexity of convex optimization
Alekh Agarwal, Martin J Wainwright, Peter Bartlett, and Pradeep Ravikumar · 2009
Earlier work this paper cites.
Large-scale machine learning with stochastic gradient descent
Léon Bottou · 2010
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Large-scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Marc’aurelio Ranzato, Andrew Senior, Paul Tucker, and Ke Yang · 2012
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
Alexander Rakhlin, Ohad Shamir, and Karthik Sridharan · 2012
Earlier work this paper cites.
Lecture 6.5-RMSProp: Divide the gradient by a running average of its recent magnitude
Tijmen Tieleman and Geoffrey Hinton · 2012
Earlier work this paper cites.
Identifying and attacking the saddle point problem in high-dimensional non-convex optimization
Yann N. Dauphin, Razvan Pascanu, Caglar Gulcehre, Kyunghyun Cho, Surya Ganguli, and Yoshua Bengio · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2015
Earlier work this paper cites.
On the ergodic convergence rates of a first-order primal-dual algorithm
Antonin Chambolle and Thomas Pock · 2016
Earlier work this paper cites.
Accelerated gradient methods for non-convex nonlinear and stochastic programming
Saeed Ghadimi and Guanghui Lan · 2016
Earlier work this paper cites.
Linear convergence of gradient and proximal-gradient methods under the Polyak-Łojasiewicz condition
Hamed Karimi, Julie Nutini, and Mark Schmidt · 2016
Earlier work this paper cites.
An overview of gradient descent optimization algorithms
Sebastian Ruder · 2016
Earlier work this paper cites.
Variants of RMSProp and AdaGrad with logarithmic regret bounds
Mahesh Chandra Mukkamala and Matthias Hein · 2017
Earlier work this paper cites.
On exponential convergence of SGD in non-convex over-parametrized learning
Raef Bassily, Mikhail Belkin, and Siyuan Ma · 2018
Cited alongside, same era.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2018
Cited alongside, same era.
On the convergence of Adam and beyond
Sashank J Reddi, Satyen Kale, and Sanjiv Kumar · 2018
Cited alongside, same era.
Adaptive methods for non-convex optimization
Manzil Zaheer, Sashank Reddi, Devendra Sachan, Satyen Kale, and Sanjiv Kumar · 2018
Cited alongside, same era.
Weighted AdaGrad with unified momentum
Fangyu Zou, Li Shen, Zequn Jie, Ju Sun, and Wei Liu · 2018
Cited alongside, same era.
On the convergence of a class of Adam-type algorithms for non-convex optimization
Convergence analysis of AdaBound with relaxed bound functions for non-convex optimization
Jinlan Liu, Jun Kong, Dongpo Xu, Miao Qi, and Yinghua Lu · 2021
Later among the works it cites.
Almost sure convergence rates for stochastic gradient descent and stochastic heavy ball
Othmane Sebbouh, Robert M Gower, and Aaron Defazio · 2021
Later among the works it cites.
Adaptivity without compromise: A momentumized, adaptive, dual averaged gradient method for stochastic optimization
Aaron Defazio and Samy Jelassi · 2022
Later among the works it cites.
A simple convergence proof of Adam and AdaGrad
Alexandre Défossez, Leon Bottou, Francis Bach, and Nicolas Usunier · 2022
Later among the works it cites.
Stochastic gradient descent with dependent data for offline reinforcement learning
Jing Dong and Xin T Tong · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Xiangyi Chen, Sijia Liu, Ruoyu Sun, and Mingyi Hong · 2019
Cited alongside, same era.
Stochastic gradient descent for non-convex learning without bounded gradient assumptions
Yunwen Lei, Ting Hu, Guiying Li, and Ke Tang · 2019
Cited alongside, same era.
Adaptive gradient methods with dynamic bound of learning rate
Liangchen Luo, Yuanhao Xiong, and Yan Liu · 2019
Cited alongside, same era.
New convergence aspects of stochastic gradient algorithms
Lam M. Nguyen, Phuong Ha Nguyen, Peter Richtárik, Katya Scheinberg, Martin Takáč, and Marten van Dijk · 2019
Cited alongside, same era.
A deep learning method for Chinese singer identification
Zebang Shen, Binbin Yong, Gaofeng Zhang, Rui Zhou, and Qingguo Zhou · 2019
Cited alongside, same era.
Non-ergodic convergence analysis of Heavy-ball algorithms
Tao Sun, Penghang Yin, Dongsheng Li, Chun Huang, Lei Guan, and Hao Jiang · 2019
Cited alongside, same era.
A sufficient condition for convergences of Adam and RMSProp
Fangyu Zou, Li Shen, Zequn Jie, Weizhong Zhang, and Wei Liu · 2019
Cited alongside, same era.
Asymptotic study of stochastic adaptive algorithms in non-convex landscape
Sébastien Gadat and Ioana Gavra · 2022
Later among the works it cites.
Revisit last-iterate convergence of mSGD under milder requirement on step size
Xingkang He, Lang Chen, Difei Cheng, Vijay Gupta, et al · 2022
Later among the works it cites.
High probability bounds for a class of nonconvex algorithms with AdaGrad stepsize
Ali Kavis, Kfir Levy, and Volkan Cevher · 2022
Later among the works it cites.
Simple and optimal stochastic gradient methods for nonsmooth nonconvex optimization
Zhize Li and Jian Li · 2022
Later among the works it cites.
On hyper-parameter selection for guaranteed convergence of RMSProp
Jinlan Liu, Dongpo Xu, Huisheng Zhang, and Danilo Mandic · 2022
Later among the works it cites.
On almost sure convergence rates of stochastic gradient methods
Jun Liu and Ye Yuan · 2022
Later among the works it cites.
SGD-r α \alpha : A real-time α \alpha -suffix averaging method for SGD with biased gradient estimates
Jianqi Luo, Jinlan Liu, Dongpo Xu, and Huisheng Zhang · 2022
Later among the works it cites.
Bohan Wang, Yushun Zhang, Huishuai Zhang, Qi Meng, Zhiming Ma, Tieyan Liu, and Wei Chen · 2022
Later among the works it cites.
Adam can converge without any modification on update rules
Yushun Zhang, Congliang Chen, Naichen Shi, Ruoyu Sun, and Zhiquan Luo · 2022
Later among the works it cites.
Better theory for SGD in the nonconvex world
Ahmed Khaled and Peter Richtárik · 2023
Closest in time.
Last-iterate convergence analysis of stochastic momentum methods for neural networks
Jinlan Liu, Dongpo Xu, Yinghua Lu, Jun Kong, and Danilo P. Mandic · 2023
Closest in time.
Convergence of AdaGrad for non-convex objectives: Simple proofs and relaxed assumptions
Bohan Wang, Huishuai Zhang, Zhiming Ma, and Wei Chen · 2023
Closest in time.
On the convergence of stochastic gradient descent with bandwidth-based step size
Xiaoyu Wang and Yaxiang Yuan · 2023
Closest in time.