Fetching the paper…
Reading the bibliography…
The adaptive stochastic gradient descent (SGD) with momentum has been widely adopted in deep learning as well as convex optimization.
Heavy-ball algorithms always escape saddle points
Tao Sun, Dongsheng Li, Zhe Quan, Hao Jiang, Shengguo Li, and Yong Dou · 1907
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
Boris T Polyak · 1964
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
Arkadi Semenovich Nemirovsky and David Borisovich Yudin · 1983
Earlier work this paper cites.
A method of solving a convex programming problem with convergence rate o ( 1 / k 2 ) o(1/k^{2})
Yu Nesterov · 1983
Earlier work this paper cites.
Convex analysis and optimization
Bertsekas Dimitri P., Nedić Angelia., and Ozdaglar Asuman E · 2003
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Martin Zinkevich · 2003
Earlier work this paper cites.
Accelerated gradient methods for stochastic optimization and online learning
Chonghai Hu, Weike Pan, and James T Kwok · 2009
Earlier work this paper cites.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Earlier work this paper cites.
Making gradient descent optimal for strongly convex stochastic optimization
Alexander Rakhlin, Ohad Shamir, and Karthik Sridharan · 2011
Earlier work this paper cites.
Optimal regularized dual averaging methods for stochastic optimization
Xi Chen, Qihang Lin, and Javier Pena · 2012
Earlier work this paper cites.
Imagenet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever, and Geoffrey E Hinton · 2012
Earlier work this paper cites.
Open problem: Is averaging needed for strongly convex stochastic gradient descent?
Ohad Shamir · 2012
Earlier work this paper cites.
Lecture 6.5-rmsprop, coursera: Neural networks for machine learning
Tijmen Tieleman and Geoffrey Hinton · 2012
Cited alongside, same era.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Cited alongside, same era.
The unusual effectiveness of averaging in gan training
Diederik P Kingma and Jimmy Ba · 2014
Cited alongside, same era.
ipiano: Inertial proximal algorithm for nonconvex optimization
Peter Ochs, Yunjin Chen, Thomas Brox, and Thomas Pock · 2014
Cited alongside, same era.
Global convergence of the heavy-ball method for convex optimization
Euhanna Ghadimi, Hamid Reza Feyzmahdavian, and Mikael Johansson · 2015
Cited alongside, same era.
Weighted adagrad with unified momentum
Fangyu Zou, Li Shen, Zequn Jie, Ju Sun, and Wei Liu · 2018
Later among the works it cites.
On the convergence of a class of adam-type algorithms for non-convex optimization
Xiangyi Chen, Sijia Liu, Ruoyu Sun, and Mingyi Hong · 2019
Later among the works it cites.
Understanding the role of momentum in stochastic gradient methods
Igor Gitman, Hunter Lang, Pengchuan Zhang, and Lin Xiao · 2019
Later among the works it cites.
Tight analyses for non-smooth stochastic gradient descent
Nicholas JA Harvey, Christopher Liaw, Yaniv Plan, and Sikander Randhawa · 2019
Later among the works it cites.
Making the last iterate of sgd information theoretically optimal
Prateek Jain, Dheeraj Nagaraj, and Praneeth Netrapalli · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Sebastian Ruder · 2016
Cited alongside, same era.
Unified convergence analysis of stochastic momentum methods for convex and non-convex optimization
Tianbao Yang, Qihang Lin, and Zhe Li · 2016
Cited alongside, same era.
Variants of rmsprop and adagrad with logarithmic regret bounds
Mahesh Chandra Mukkamala and Matthias Hein · 2017
Cited alongside, same era.
The marginal value of adaptive gradient methods in machine learning
Ashia C Wilson, Rebecca Roelofs, Mitchell Stern, Nati Srebro, and Benjamin Recht · 2017
Cited alongside, same era.
Introductory lectures on stochastic optimization
John C Duchi · 2018
Cited alongside, same era.
On the convergence of adam and beyond
Sashank J Reddi, Satyen Kale, and Sanjiv Kumar · 2018
Cited alongside, same era.
Non-ergodic convergence analysis of heavy-ball algorithms
Tao Sun, Penghang Yin, Dongsheng Li, Chun Huang, L. Guan, and Hao Jiang
Cited in the paper.
Antonio Orvieto, Jonas Köhler, and A. Lucchi · 2019
Later among the works it cites.
A new regret analysis for adam-type algorithms
Ahmet Alacaoglu, Yura Malitsky, Panayotis Mertikopoulos, and Volkan Cevher · 2020
Later among the works it cites.
On the convergence of adam and adagrad
Alexandre Défossez, L. Bottou, Francis R. Bach, and Nicolas Usunier · 2020
Later among the works it cites.
Accelerating sgd with momentum for over-parameterized learning
Chaoyue Liu and Mikhail Belkin · 2020
Later among the works it cites.
On the convergence of the stochastic heavy ball method
Othmane Sebbouh, Robert Mansel Gower, and Aaron Defazio · 2020
Later among the works it cites.
Sadam: A variant of adam for strongly convex functions
Guanghui Wang, Shiyin Lu, Weiwei Tu, and Lijun Zhang · 2020
Later among the works it cites.