Fetching the paper…
Reading the bibliography…
Minimax optimal convergence rates for classes of stochastic convex optimization problems are well characterized, where the majority of results utilize iterate averaged stochastic gradient descent (SGD) with polynomially decaying step sizes.
Making the last iterate of sgd information theoretically optimal
P. Jain, D. Nagaraj, and P. Netrapalli · 1904
Earlier work this paper cites.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Adaptive switching circuits
B. Widrow and M. E. Hoff · 1960
Earlier work this paper cites.
A learning method for system identification
J.-I. Nagumo and A. Noda · 1967
Earlier work this paper cites.
On Optimal Estimation Methods Using Stochastic Approximation Procedures
D. Anbar · 1971
Earlier work this paper cites.
Channel identification for high speed digital communications
J. G. Proakis · 1974
Earlier work this paper cites.
On the convergence rates of subgradient optimization methods
J. L. Goffin · 1977
Earlier work this paper cites.
Stochastic Approximation Methods for Constrained and Unconstrained Systems
H. J. Kushner and D. S. Clark · 1978
Earlier work this paper cites.
Problem Complexity and Method Efficiency in Optimization
A. S. Nemirovsky and D. B. Yudin · 1983
Earlier work this paper cites.
A method for unconstrained convex minimization problem with the rate of convergence O ( 1 / k 2 ) {O}(1/k^{2})
Y. E. Nesterov · 1983
Earlier work this paper cites.
Adaptive Signal Processing
B. Widrow and S. D. Stearns · 1985
Earlier work this paper cites.
Efficient estimations from a slowly convergent robbins-monro process
D. Ruppert · 1988
Earlier work this paper cites.
Adaptive Algorithms and Stochastic Approximations
A. Benveniste, M. Metivier, and P. Priouret · 1990
Earlier work this paper cites.
Analysis of the momentum lms algorithm
S. Roy and J. J. Shynk · 1990
Earlier work this paper cites.
Stochastic Approximation and Optimization of Random Systems
L. Ljung, G. Pflug, and H. Walk · 1992
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
B. T. Polyak and A. B. Juditsky · 1992
Earlier work this paper cites.
Theory of Point Estimation
E. L. Lehmann and G. Casella · 1998
Earlier work this paper cites.
Analysis of momentum adaptive filtering algorithms
R. Sharma, W. A. Sethares, and J. A. Bucklew · 1998
Earlier work this paper cites.
Stochastic approximation algorithms: overview and recent trends
B. Bharath and V. S. Borkar · 1999
Earlier work this paper cites.
Asymptotic statistics , volume 3
A. W. Van der Vaart · 2000
Cited alongside, same era.
S. Lacoste-Julien, M. W. Schmidt, and F. R. Bach · 2002
Cited alongside, same era.
Stochastic approximation and recursive algorithms and applications
H. J. Kushner and G. Yin · 2003
Cited alongside, same era.
Stochastic approximation: invited paper, 2003
T. L. Lai · 2003
Cited alongside, same era.
The tradeoffs of large scale learning
L. Bottou and O. Bousquet · 2007
Cited alongside, same era.
Stochastic approximation
V. Borkar · 2008
Cited alongside, same era.
Theory of convex optimization for machine learning
S. Bubeck · 2014
Later among the works it cites.
Beyond the regret minimization barrier: optimal algorithms for stochastic strongly-convex optimization
E. Hazan and S. Kale · 2014
Later among the works it cites.
Averaged least-mean-squares: Bias-variance trade-offs and optimal sampling distributions
A. Défossez and F. R. Bach · 2015
Later among the works it cites.
Non-parametric stochastic approximation with large step sizes
A. Dieuleveut and F. R. Bach · 2015
Later among the works it cites.
Competing with the empirical risk minimizer in a single pass
R. Frostig, R. Ge, S. M. Kakade, and A. Sidford · 2015
Later among the works it cites.
Harder, better, faster, stronger convergence rates for least-squares regression
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
F. R. Bach and E. Moulines · 2011
Cited alongside, same era.
Adaptive subgradient methods for online learning and stochastic optimization
J. C. Duchi, E. Hazan, and Y. Singer · 2011
Cited alongside, same era.
Information-based complexity, feedback and dynamics in convex programming
M. Raginsky and A. Rakhlin · 2011
Cited alongside, same era.
Information-theoretic lower bounds on the oracle complexity of stochastic convex optimization
A. Agarwal, P. L. Bartlett, P. Ravikumar, and M. J. Wainwright · 2012
Cited alongside, same era.
Optimal distributed online prediction using mini-batches
O. Dekel, R. Gilad-Bachrach, O. Shamir, and L. Xiao · 2012
Cited alongside, same era.
Optimal stochastic approximation algorithms for strongly convex stochastic composite optimization i: A generic algorithmic framework
S. Ghadimi and G. Lan · 2012
Cited alongside, same era.
A. Dieuleveut, N. Flammarion, and F. R. Bach · 2016
Later among the works it cites.
Parallelizing stochastic approximation through mini-batching and tail-averaging
P. Jain, S. M. Kakade, R. Kidambi, P. Netrapalli, and A. Sidford · 2016
Later among the works it cites.
Stochastic gradient descent, weighted sampling, and the randomized kaczmarz algorithm
D. Needell, N. Srebro, and R. Ward · 2016
Later among the works it cites.
Accelerate stochastic subgradient method by leveraging local error bound
Y. Xu, Q. Lin, and T. Yang · 2016
Later among the works it cites.
Hyperband: A novel bandit-based approach to hyperparameter optimization
L. Li, K. Jamieson, G. DeSalvo, A. Rostamizadeh, and A. Talwalkar · 2017
Later among the works it cites.
How to make the gradients small stochastically
Z. Allen-Zhu · 2018
Later among the works it cites.
Tight analyses for non-smooth stochastic gradient descent
N. J. A. Harvey, C. Liaw, Y. Plan, and S. Randhawa · 2018
Later among the works it cites.
Iterate averaging as regularization for stochastic gradient descent
G. Neu and L. Rosasco · 2018
Later among the works it cites.
Why does stagewise training accelerate convergence of testing error over sgd?
T. Yang, Y. Y. 0006, Z. Yuan, and R. Jin · 2018
Later among the works it cites.
A universally optimal multistage accelerated stochastic gradient method
N. S. Aybat, A. Fallah, M. Gürbüzbalaban, and A. E. Ozdaglar · 2019
Closest in time.
Robust stochastic optimization with the proximal point method
D. Davis and D. Drusvyatskiy · 2019
Closest in time.
Stochastic algorithms with geometric step decay converge linearly on sharp functions
D. Davis, D. Drusvyatskiy, and V. Charisopoulos · 2019
Closest in time.
A generic acceleration framework for stochastic composite optimization
A. Kulunchakov and J. Mairal · 2019
Closest in time.