Fetching the paper…
Reading the bibliography…
This work provides a simplified proof of the statistical minimax optimality of (iterate averaged) stochastic gradient descent (SGD), for the special case of least squares.
Stochastic Approximation Methods for Constrained and Unconstrained Systems
Harold J. Kushner and Dean S. Clark · 1978
Earlier work this paper cites.
Efficient estimations from a slowly convergent robbins-monro process
David Ruppert · 1988
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
Boris T. Polyak and Anatoli B. Juditsky · 1992
Earlier work this paper cites.
Theory of Point Estimation
Erich L. Lehmann and George Casella · 1998
Earlier work this paper cites.
Aad W. van der Vaart · 2000
Cited alongside, same era.
Adaptivity of averaged stochastic gradient descent to local strong convexity for logistic regression
Francis R. Bach · 2014
Cited alongside, same era.
Averaged least-mean-squares: Bias-variance trade-offs and optimal sampling distributions
Alexandre Défossez and Francis R. Bach · 2015
Cited alongside, same era.
Non-parametric stochastic approximation with large step sizes
Aymeric Dieuleveut and Francis R. Bach · 2015
Later among the works it cites.
Competing with the empirical risk minimizer in a single pass
Roy Frostig, Rong Ge, Sham M. Kakade, and Aaron Sidford · 2015
Later among the works it cites.
Parallelizing stochastic approximation through mini-batching and tail-averaging
Prateek Jain, Sham M. Kakade, Rahul Kidambi, Praneeth Netrapalli, and Aaron Sidford · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…