Fetching the paper…
Reading the bibliography…
Stochastic gradient descent (SGD) is one of the most widely used algorithms for large scale optimization problems.
Problem complexity and method efficiency in optimization
Arkadii Semenovich Nemirovsky and David Borisovich Yudin · 1983
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Boris T Polyak and Anatoli B Juditsky · 1992
Earlier work this paper cites.
Online convex programming and generalized infinitesimal gradient ascent
Martin Zinkevich · 2003
Earlier work this paper cites.
On the generalization ability of on-line learning algorithms
Nicolo Cesa-Bianchi, Alex Conconi, and Claudio Gentile · 2004
Earlier work this paper cites.
Solving large scale linear prediction problems using stochastic gradient descent algorithms
Tong Zhang · 2004
Earlier work this paper cites.
On the generalization ability of online strongly convex programming algorithms
Sham M Kakade and Ambuj Tewari · 2009
Earlier work this paper cites.
Pegasos: Primal estimated sub-gradient solver for svm
Shai Shalev-Shwartz, Yoram Singer, Nathan Srebro, and Andrew Cotter · 2011
Cited alongside, same era.
Open problem: Is averaging needed for strongly convex stochastic gradient descent?
Ohad Shamir · 2012
Cited alongside, same era.
Making gradient descent optimal for strongly convex stochastic optimization
Alexander Rakhlin, Ohad Shamir, Karthik Sridharan, et al · 2012
Cited alongside, same era.
Simon Lacoste-Julien, Mark Schmidt, and Francis Bach · 2012
Cited alongside, same era.
Stochastic gradient descent for non-smooth optimization: Convergence results and optimal averaging schemes
Ohad Shamir and Tong Zhang · 2013
Cited alongside, same era.
Beyond the regret minimization barrier: optimal algorithms for stochastic strongly-convex optimization
Elad Hazan and Satyen Kale · 2014
Later among the works it cites.
Deep learning
Yann LeCun, Yoshua Bengio, and Geoffrey Hinton · 2015
Later among the works it cites.
Convex optimization: Algorithms and complexity
Sébastien Bubeck · 2015
Later among the works it cites.
Extremely large minibatch sgd: Training resnet-50 on imagenet in 15 minutes
Takuya Akiba, Shuji Suzuki, and Keisuke Fukuda · 2017
Later among the works it cites.
Tight analyses for non-smooth stochastic gradient descent
Nicholas JA Harvey, Christopher Liaw, Yaniv Plan, and Sikander Randhawa · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…