Fetching the paper…
Reading the bibliography…
We study the convergence of Stochastic Gradient Descent (SGD) for strongly convex objective functions.
A stochastic approximation method
Herbert Robbins and Sutton Monro · 1951
Earlier work this paper cites.
Convergence of estimates under dimensionality restrictions
Lucien LeCam et al · 1973
Earlier work this paper cites.
Elements of Information Theory
Thomas M. Cover and Joy A. Thomas · 1991
Earlier work this paper cites.
Assouad, Fano, and Le Cam
Bin Yu · 1997
Earlier work this paper cites.
Nonlinear Programming
D.P. Bertsekas · 1999
Earlier work this paper cites.
Convex Optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Introductory lectures on convex optimization : a basic course
Yurii Nesterov · 2004
Earlier work this paper cites.
Numerical Optimization
Jorge Nocedal and Stephen J. Wright · 2006
Earlier work this paper cites.
Information-theoretic lower bounds on the oracle complexity of convex optimization
Alekh Agarwal, Martin J Wainwright, Peter L Bartlett, and Pradeep K Ravikumar · 2009
Cited alongside, same era.
The Elements of Statistical Learning: Data Mining, Inference, and Prediction
Trevor Hastie, Robert Tibshirani, and Jerome Friedman · 2009
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
John Duchi, Elad Hazan, and Yoram Singer · 2011
Cited alongside, same era.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Eric Moulines and Francis R Bach · 2011
Cited alongside, same era.
Information-Based Complexity, Feedback and Dynamics in Convex Programming
Maxim Raginsky and Alexander Rakhlin · 2011
Cited alongside, same era.
A stochastic gradient method with an exponential convergence rate for finite training sets
On the importance of initialization and momentum in deep learning
Ilya Sutskever, James Martens, George Dahl, and Geoffrey Hinton · 2013
Later among the works it cites.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
Aaron Defazio, Francis Bach, and Simon Lacoste-Julien · 2014
Later among the works it cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba · 2014
Later among the works it cites.
Optimization methods for large-scale machine learning
Léon Bottou, Frank E Curtis, and Jorge Nocedal · 2016
Later among the works it cites.
SARAH: A novel method for machine learning problems using stochastic recursive gradient
Lam M. Nguyen, Jie Liu, Katya Scheinberg, and Martin Takáč · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nicolas Le Roux, Mark Schmidt, and Francis Bach · 2012
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Rie Johnson and Tong Zhang · 2013
Cited alongside, same era.
Fast convergence of stochastic gradient descent under a strong growth condition
Mark Schmidt and Nicolas Le Roux · 2013
Cited alongside, same era.
Problem complexity and method efficiency in optimization /
Arkadii Semenovich. Nemirovsky A.S. and D. B. IUdin
Cited in the paper.
New convergence aspects of stochastic gradient algorithms
Lam Nguyen, Phuong Ha Nguyen, Peter Richtarik, Katya Scheinberg, Martin Takac, and Marten van Dijk
Cited in the paper.
SGD and hogwild! Convergence without the bounded gradients assumption
Lam Nguyen, Phuong Ha Nguyen, Marten van Dijk, Peter Richtarik, Katya Scheinberg, and Martin Takac
Cited in the paper.
Improved asynchronous parallel optimization analysis for stochastic incremental methods
Rémi Leblond, Fabian Pederegosa, and Simon Lacoste-Julien · 2018
Closest in time.
A Sufficient Condition for Convergences of Adam and RMSProp
Fangyu Zou, Li Shen, Zequn Jie, Weizhong Zhang, and Wei Liu · 2018
Closest in time.
Sgd: General analysis and improved rates
Robert Mansel Gower, Nicolas Loizou, Xun Qian, Alibek Sailanbayev, Egor Shulgin, and Peter Richtarik · 2019
Closest in time.