Fetching the paper…
Reading the bibliography…
We study stochastic gradient descent (SGD) and the stochastic heavy ball method (SHB, otherwise known as the momentum method) for the general stochastic approximation problem.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
B. T. Polyak · 1964
Earlier work this paper cites.
A convergence theorem for nonnegative almost supermartingales and some applications
Herbert Robbins and David Siegmund · 1971
Earlier work this paper cites.
Introduction to optimization, translations series in mathematics and engineering
BT Polyak · 1987
Earlier work this paper cites.
Gradient convergence in gradient methods with errors
Dimitri P Bertsekas and John N Tsitsiklis · 2000
Earlier work this paper cites.
Stochastic learning
Leon Bottou · 2003
Earlier work this paper cites.
Sequential quadratic programming
Jorge Nocedal and Stephen J Wright · 2006
Earlier work this paper cites.
A fast iterative shrinkage-thresholding algorithm with application to wavelet-based image deblurring
Amir Beck and Marc Teboulle · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Arkadi Nemirovski, Anatoli B. Juditsky, Guanghui Lan, and Alexander Shapiro · 2009
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Francis Bach and Eric Moulines · 2011
Earlier work this paper cites.
Convex Analysis and Monotone Operator Theory in Hilbert Spaces
Heinz H. Bauschke and Patrick L. Combettes · 2011
Earlier work this paper cites.
Stochastic first- and zeroth-order methods for nonconvex stochastic programming
Saeed Ghadimi and Guanghui Lan · 2013
Earlier work this paper cites.
Introductory lectures on convex optimization: A basic course , volume 87
Yurii Nesterov · 2013
Earlier work this paper cites.
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Earlier work this paper cites.
On the convergence of the iterates of the "fast iterative shrinkage/thresholding algorithm"
Antonin Chambolle and Charles Dossal · 2015
Cited alongside, same era.
Global convergence of the heavy-ball method for convex optimization
Euhanna Ghadimi, Hamid Feyzmahdavian, and Mikael Johansson · 2015
Cited alongside, same era.
The rate of convergence of nesterov’s accelerated forward-backward method is actually faster than 1/k 2 {}^{\mbox{2}}
Hédy Attouch and Juan Peypouquet · 2016
Cited alongside, same era.
Antoine Godichon-Baggioni · 2016
Cited alongside, same era.
Unified convergence analysis of stochastic momentum methods for convex and non-convex optimization
Tianbao Yang, Qihang Lin, and Zhe Li · 2016
Optimal mini-batch and step sizes for saga
Nidham Gazagnadou, Robert Mansel Gower, and Joseph Salmon · 2019
Later among the works it cites.
SGD: General Analysis and Improved Rates
Robert Mansel Gower, Nicolas Loizou, Xun Qian, Alibek Sailanbayev, Egor Shulgin, and Peter Richtárik · 2019
Later among the works it cites.
Making the last iterate of sgd information theoretically optimal
Prateek Jain, Dheeraj Nagaraj, and Praneeth Netrapalli · 2019
Later among the works it cites.
First-order algorithms converge faster than o ( 1 / k ) o(1/k) on convex problems
Ching-Pei Lee and Stephen Wright · 2019
Later among the works it cites.
On the convergence of stochastic gradient descent with adaptive stepsizes
Xiaoyu Li and Francesco Orabona · 2019
Later among the works it cites.
The role of memory in stochastic optimization
Antonio Orvieto, Jonas Köhler, and Aurélien Lucchi · 2019
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
An optimal randomized incremental gradient method
Guanghui Lan and Yi Zhou · 2017
Cited alongside, same era.
Stochastic mirror descent in variationally coherent optimization problems
Zhengyuan Zhou, Panayotis Mertikopoulos, Nicholas Bambos, Stephen P. Boyd, and Peter W. Glynn · 2017
Cited alongside, same era.
Stochastic heavy ball
Sébastien Gadat, Fabien Panloup, and Sofiane Saadane · 2018
Cited alongside, same era.
On the insufficiency of existing momentum schemes for stochastic optimization
Rahul Kidambi, Praneeth Netrapalli, Prateek Jain, and Sham M. Kakade · 2018
Cited alongside, same era.
Nicolas Loizou and Peter Richtárik · 2018
Cited alongside, same era.
SGD and hogwild! convergence without the bounded gradients assumption
Lam M. Nguyen, Phuong Ha Nguyen, Marten van Dijk, Peter Richtárik, Katya Scheinberg, and Martin Takác · 2018
Cited alongside, same era.
Accelerated linear convergence of stochastic momentum methods in wasserstein distances
Bugra Can, Mert Gürbüzbalaban, and Lingjiong Zhu · 2019
Cited alongside, same era.
Later among the works it cites.
Fast and faster convergence of SGD for over-parameterized models and an accelerated perceptron
Sharan Vaswani, Francis Bach, and Mark Schmidt · 2019
Later among the works it cites.
Better theory for SGD in the nonconvex world
Ahmed Khaled and Peter Richtárik · 2020
Closest in time.
Stochastic polyak step-size for sgd: An adaptive learning rate for fast convergence
Nicolas Loizou, Sharan Vaswani, Issam Laradji, and Simon Lacoste-Julien · 2020
Closest in time.
On the almost sure convergence of stochastic gradient descent in non-convex problems
Panayotis Mertikopoulos, Nadav Hallak, Ali Kavis, and Volkan Cevher · 2020
Closest in time.
Adaptive gradient methods converge faster with over-parameterization (and you can do a line-search)
Sharan Vaswani, Frederik Kunstner, Issam Laradji, Si Yi Meng, Mark Schmidt, and Simon Lacoste-Julien · 2020
Closest in time.
Almost sure convergence of sgd on smooth non-convex functions
Francesco Orabona · 2021
Closest in time.
Last iterate of sgd converges (even in bounded domains)
Francesco Orabona · 2021
Closest in time.