Fetching the paper…
Reading the bibliography…
This work characterizes the benefits of averaging schemes widely used in conjunction with stochastic gradient descent (SGD).
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Some methods of speeding up the convergence of iteration methods
B. T. Polyak · 1964
Earlier work this paper cites.
On Optimal Estimation Methods Using Stochastic Approximation Procedures
D. Anbar · 1971
Earlier work this paper cites.
Asymptotically efficient stochastic approximation; the RM case
V. Fabian · 1973
Earlier work this paper cites.
Stochastic Approximation Methods for Constrained and Unconstrained Systems
H. J. Kushner and D. S. Clark · 1978
Earlier work this paper cites.
Problem Complexity and Method Efficiency in Optimization
A. S. Nemirovsky and D. B. Yudin · 1983
Earlier work this paper cites.
A method for unconstrained convex minimization problem with the rate of convergence O ( 1 / k 2 ) {O}(1/k^{2})
Y. E. Nesterov · 1983
Earlier work this paper cites.
Asymptotic properties of distributed and communicating stochastic approximation algorithms
H. J. Kushner and G. Yin · 1987
Earlier work this paper cites.
Efficient estimations from a slowly convergent robbins-monro process
D. Ruppert · 1988
Earlier work this paper cites.
Acceleration of stochastic approximation by averaging
B. T. Polyak and A. B. Juditsky · 1992
Earlier work this paper cites.
Theory of Point Estimation
E. L. Lehmann and G. Casella · 1998
Earlier work this paper cites.
Asymptotic Statistics
A. W. van der Vaart · 2000
Earlier work this paper cites.
Stochastic approximation and recursive algorithms and applications
H. J. Kushner and G. Yin · 2003
Earlier work this paper cites.
Positive Definite Matrices
R. Bhatia · 2007
Earlier work this paper cites.
The tradeoffs of large scale learning
L. Bottou and O. Bousquet · 2007
Earlier work this paper cites.
Efficient large-scale distributed training of conditional maximum entropy models
G. Mann, R. T. McDonald, M. Mohri, N. Silberman, and D. Walker · 2009
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
F. Bach and E. Moulines · 2011
Cited alongside, same era.
Parallel coordinate descent for l1-regularized loss minimization
J. K. Bradley, A. Kyrola, D. Bickson, and C. Guestrin · 2011
Cited alongside, same era.
Better mini-batch algorithms via accelerated gradient methods
A. Cotter, O. Shamir, N. Srebro, and K. Sridharan · 2011
Cited alongside, same era.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
F. Niu, B. Recht, C. Re, and S. J. Wright · 2011
Cited alongside, same era.
Parallelized stochastic gradient descent
M. A. Zinkevich, A. Smola, M. Weimer, and L. Li · 2011
Cited alongside, same era.
Information-theoretic lower bounds on the oracle complexity of stochastic convex optimization
On the optimality of averaging in distributed statistical learning
J. Rosenblatt and B. Nadler · 2014
Later among the works it cites.
Averaged least-mean-squares: Bias-variance trade-offs and optimal sampling distributions
A. Défossez and F. R. Bach · 2015
Later among the works it cites.
Non-parametric stochastic approximation with large step sizes
A. Dieuleveut and F. Bach · 2015
Later among the works it cites.
Asynchronous stochastic convex optimization
J. C. Duchi, S. Chaturapruek, and C. Ré · 2015
Later among the works it cites.
A universal catalyst for first-order optimization
H. Lin, J. Mairal, and Z. Harchaoui · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A. Agarwal, P. L. Bartlett, P. Ravikumar, and M. J. Wainwright · 2012
Cited alongside, same era.
Optimal distributed online prediction using mini-batches
O. Dekel, R. Gilad-Bachrach, O. Shamir, and L. Xiao · 2012
Cited alongside, same era.
A stochastic gradient method with an exponential convergence rate for strongly-convex optimization with finite training sets
N. L. Roux, M. Schmidt, and F. R. Bach · 2012
Cited alongside, same era.
Stochastic dual coordinate ascent methods for regularized loss minimization
S. Shalev-Shwartz and T. Zhang · 2012
Cited alongside, same era.
Non-strongly-convex smooth stochastic approximation with convergence rate O(1/n)
F. R. Bach and E. Moulines · 2013
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
R. Johnson and T. Zhang · 2013
Cited alongside, same era.
Mini-batch primal and dual methods for SVMs
M. Takác, A. S. Bijral, P. Richtárik, and N. Srebro · 2013
Cited alongside, same era.
M. Takác, P. Richtárik, and N. Srebro · 2015
Later among the works it cites.
Disco: Distributed optimization for self-concordant empirical loss
Y. Zhang and L. Xiao · 2015
Later among the works it cites.
Divide and conquer ridge regression: A distributed algorithm with minimax optimal rates
Y. Zhang, J. C. Duchi, and M. Wainwright · 2015
Later among the works it cites.
Katyusha: The first direct acceleration of stochastic gradient methods
Z. Allen-Zhu · 2016
Closest in time.
Optimization methods for large-scale machine learning
L. Bottou, F. E. Curtis, and J. Nocedal · 2016
Closest in time.
A simple practical accelerated method for finite sums
A. Defazio · 2016
Closest in time.
Stochastic gradient descent, weighted sampling, and the randomized kaczmarz algorithm
D. Needell, N. Srebro, and R. Ward · 2016
Closest in time.
Accurate, large minibatch sgd: training imagenet in 1 hour
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Closest in time.
Regularizing and optimizing lstm language models
S. Merity, N. S. Keskar, and R. Socher · 2017
Closest in time.
Don’t decay the learning rate, increase the batch size
S. L. Smith, P.-J. Kindermans, and Q. V. Le · 2017
Closest in time.