Fetching the paper…
Reading the bibliography…
Stochastic gradient descent (SGD) is the optimization algorithm of choice in many machine learning applications such as regularized empirical risk minimization and training deep neural networks.
A stochastic approximation method
Robbins, Herbert and Monro, Sutton · 1951
Earlier work this paper cites.
Introductory lectures on convex optimization : a basic course
Nesterov, Yurii · 2004
Earlier work this paper cites.
Numerical Optimization
Nocedal, Jorge and Wright, Stephen J · 2006
Earlier work this paper cites.
Pegasos: Primal estimated sub-gradient solver for svm
Shalev-Shwartz, Shai, Singer, Yoram, and Srebro, Nathan · 2007
Earlier work this paper cites.
The Elements of Statistical Learning: Data Mining, Inference, and Prediction
Hastie, Trevor, Tibshirani, Robert, and Friedman, Jerome · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
Nemirovski, A., Juditsky, A., Lan, G., and Shapiro, A · 2009
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
Moulines, Eric and Bach, Francis R · 2011
Earlier work this paper cites.
Hogwild!: A Lock-Free Approach to Parallelizing Stochastic Gradient Descent
Recht, Benjamin, Re, Christopher, Wright, Stephen, and Niu, Feng · 2011
Cited alongside, same era.
A stochastic gradient method with an exponential convergence rate for finite training sets
Le Roux, Nicolas, Schmidt, Mark, and Bach, Francis · 2012
Cited alongside, same era.
Making gradient descent optimal for strongly convex stochastic optimization
Rakhlin, Alexander, Shamir, Ohad, and Sridharan, Karthik · 2012
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
Johnson, Rie and Zhang, Tong · 2013
Cited alongside, same era.
SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives
Defazio, Aaron, Bach, Francis, and Lacoste-Julien, Simon · 2014
Cited alongside, same era.
Incremental gradient, subgradient, and proximal methods for convex optimization: A survey, 2015
Bertsekas, Dimitri P · 2015
Later among the works it cites.
Taming the wild: A unified analysis of hogwild-style algorithms
De Sa, Christopher M, Zhang, Ce, Olukotun, Kunle, and Ré, Christopher · 2015
Later among the works it cites.
Perturbed Iterate Analysis for Asynchronous Stochastic Optimization
Mania, Horia, Pan, Xinghao, Papailiopoulos, Dimitris, Recht, Benjamin, Ramchandran, Kannan, and Jordan, Michael I · 2015
Later among the works it cites.
Optimization methods for large-scale machine learning
Bottou, Léon, Curtis, Frank E, and Nocedal, Jorge · 2016
Later among the works it cites.
SARAH: A novel method for machine learning problems using stochastic recursive gradient
Nguyen, Lam, Liu, Jie, Scheinberg, Katya, and Takáč, Martin · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hazan, Elad and Kale, Satyen · 2014
Cited alongside, same era.
Improved asynchronous parallel optimization analysis for stochastic incremental methods
Leblond, Remi, Pedregosa, Fabian, and Lacoste-Julien, Simon · 2018
Closest in time.