Fetching the paper…
Reading the bibliography…
There is widespread sentiment that it is not possible to effectively utilize fast gradient methods (e.g.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Methods of conjuate gradients for solving linear systems
M. R. Hestenes and E. Stiefel · 1952
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
B. T. Polyak · 1964
Earlier work this paper cites.
On Optimal Estimation Methods Using Stochastic Approximation Procedures
D. Anbar · 1971
Earlier work this paper cites.
The computation of eigenvalues and eigenvectors of very large sparse matrices
C. C. Paige · 1971
Earlier work this paper cites.
Asymptotically efficient stochastic approximation; the RM case
V. Fabian · 1973
Earlier work this paper cites.
Channel identification for high speed digital communications
J. G. Proakis · 1974
Earlier work this paper cites.
Stochastic Approximation Methods for Constrained and Unconstrained Systems
H. J. Kushner and D. S. Clark · 1978
Earlier work this paper cites.
Problem Complexity and Method Efficiency in Optimization
A. S. Nemirovsky and D. B. Yudin · 1983
Earlier work this paper cites.
A method for unconstrained convex minimization problem with the rate of convergence O ( 1 / k 2 ) {O}(1/k^{2})
Y. E. Nesterov · 1983
Earlier work this paper cites.
Adaptive Signal Processing
B. Widrow and S. D. Stearns · 1985
Earlier work this paper cites.
Introduction to Optimization
B. T. Polyak · 1987
Earlier work this paper cites.
Efficient estimations from a slowly convergent robbins-monro process
D. Ruppert · 1988
Earlier work this paper cites.
Behavior of slightly perturbed lanczos and conjugate-gradient recurrences
A. Greenbaum · 1989
Earlier work this paper cites.
Analysis of the momentum lms algorithm
S. Roy and J. J. Shynk · 1990
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
B. T. Polyak and A. B. Juditsky · 1992
Earlier work this paper cites.
Theory of Point Estimation
E. L. Lehmann and G. Casella · 1998
Cited alongside, same era.
Analysis of momentum adaptive filtering algorithms
R. Sharma, W. A. Sethares, and J. A. Bucklew · 1998
Cited alongside, same era.
Asymptotic Statistics
A. W. van der Vaart · 2000
Cited alongside, same era.
Stochastic approximation and recursive algorithms and applications
H. J. Kushner and G. Yin · 2003
Cited alongside, same era.
Introductory lectures on convex optimization: A basic course , volume 87 of Applied Optimization
Y. E. Nesterov · 2004
Cited alongside, same era.
The tradeoffs of large scale learning
L. Bottou and O. Bousquet · 2007
Cited alongside, same era.
Smooth optimization with approximate gradient
Optimal stochastic approximation algorithms for strongly convex stochastic composite optimization, ii: shrinking procedures and optimal algorithms
S. Ghadimi and G. Lan · 2013
Later among the works it cites.
Adaptivity of averaged stochastic gradient descent to local strong convexity for logistic regression
F. R. Bach · 2014
Later among the works it cites.
First-order methods of smooth convex optimization with inexact oracle
O. Devolder, F. Glineur, and Y. E. Nesterov · 2014
Later among the works it cites.
Random design analysis of ridge regression
D. J. Hsu, S. M. Kakade, and T. Zhang · 2014
Later among the works it cites.
Accelerated proximal stochastic dual coordinate ascent for regularized loss minimization
S. Shalev-Shwartz and T. Zhang · 2014
Later among the works it cites.
Averaged least-mean-squares: Bias-variance trade-offs and optimal sampling distributions
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. d’Aspremont · 2008
Cited alongside, same era.
An optimal method for stochastic composite optimization
G. Lan · 2008
Cited alongside, same era.
Accelerated gradient methods for stochastic optimization and online learning
C. Hu, J. T. Kwok, and W. Pan · 2009
Cited alongside, same era.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
F. R. Bach and E. Moulines · 2011
Cited alongside, same era.
Information-based complexity, feedback and dynamics in convex programming
M. Raginsky and A. Rakhlin · 2011
Cited alongside, same era.
Information-theoretic lower bounds on the oracle complexity of stochastic convex optimization
A. Agarwal, P. L. Bartlett, P. Ravikumar, and M. J. Wainwright · 2012
Cited alongside, same era.
A. Défossez and F. R. Bach · 2015
Later among the works it cites.
Non-parametric stochastic approximation with large step sizes
A. Dieuleveut and F. R. Bach · 2015
Later among the works it cites.
An optimal randomized incremental gradient method
G. Lan and Y. Zhou · 2015
Later among the works it cites.
A universal catalyst for first-order optimization
H. Lin, J. Mairal, and Z. Harchaoui · 2015
Later among the works it cites.
Katyusha: The first direct acceleration of stochastic gradient methods
Z. Allen-Zhu · 2016
Later among the works it cites.
Harder, better, faster, stronger convergence rates for least-squares regression
A. Dieuleveut, N. Flammarion, and F. R. Bach · 2016
Later among the works it cites.
Parallelizing stochastic approximation through mini-batching and tail-averaging
P. Jain, S. M. Kakade, R. Kidambi, P. Netrapalli, and A. Sidford · 2016
Later among the works it cites.
Stochastic gradient descent, weighted sampling, and the randomized kaczmarz algorithm
D. Needell, N. Srebro, and R. Ward · 2016
Later among the works it cites.
A lyapunov analysis of momentum methods in optimization
A. C. Wilson, B. Recht, and M. I. Jordan · 2016
Later among the works it cites.
Tight complexity bounds for optimizing composite objectives
B. Woodworth and N. Srebro · 2016
Later among the works it cites.
On the influence of momentum acceleration on online learning
K. Yuan, B. Ying, and A. H. Sayed · 2016
Later among the works it cites.