Fetching the paper…
Reading the bibliography…
In empirical risk optimization, it has been observed that stochastic gradient implementations that rely on random reshuffling of the data achieve better performance than implementations that rely on sampling the data uniformly.
“A generalization of sampling without replacement from a finite universe,”
D. G. Horvitz and D. J. Thompson, · 1952
Earlier work this paper cites.
Introduction to Optimization
B. T. Polyak, · 1987
Earlier work this paper cites.
Parallel and Distributed Computation: Numerical Methods
D. P. Bertsekas and J. N. Tsitsiklis, · 1989
Earlier work this paper cites.
“Acceleration of stochastic approximation by averaging,”
B. T. Polyak and A. B. Juditsky, · 1992
Earlier work this paper cites.
“A new class of incremental gradient methods for least squares problems,”
D. P. Bertsekas, · 1997
Earlier work this paper cites.
Convex Analysis and Optimization
D. P. Bertsekas, A. Nedi, and A. E Ozdaglar, · 2003
Earlier work this paper cites.
R. A. Horn and C. R. Johnson, · 2003
Earlier work this paper cites.
“Solving large scale linear prediction problems using stochastic gradient descent algorithms,”
T. Zhang, · 2004
Earlier work this paper cites.
Pattern Recognition and Machine Learning
C. M. Bishop, · 2006
Earlier work this paper cites.
“The tradeoffs of large scale learning,”
O. Bousquet and L. Bottou, · 2008
Earlier work this paper cites.
Pattern Recognition
S. Theodoridis and K. Koutroumbas, · 2008
Earlier work this paper cites.
“Curiously fast convergence of some stochastic gradient descent algorithms,”
L. Bottou, · 2009
Earlier work this paper cites.
“Information-theoretic lower bounds on the oracle complexity of convex optimization,”
A. Agarwal, M. J. Wainwright, P. L. Bartlett, and P. K. Ravikumar, · 2009
Cited alongside, same era.
The Elements of Statistical Learning
T. Hastie, R. Tibshirani, and J. Friedman, · 2009
Cited alongside, same era.
“Large-scale machine learning with stochastic gradient descent,”
L. Bottou, · 2010
Cited alongside, same era.
Probability: Theory and Examples
R. Durrett, · 2010
Cited alongside, same era.
“Non-asymptotic analysis of stochastic approximation algorithms for machine learning,”
E. Moulines and F. R. Bach, · 2011
Cited alongside, same era.
“Toward a noncommutative arithmetic-geometric mean inequality: Conjectures, case-studies, and consequences,”
B. Recht and C. Ré, · 2012
Cited alongside, same era.
“Adaptation, learning, and optimization over networks,”
A. H. Sayed, · 2014
Later among the works it cites.
“Adaptive networks,”
A. H. Sayed, · 2014
Later among the works it cites.
“Why random reshuffling beats stochastic gradient descent,”
M. Gürbüzbalaban, A. Ozdaglar, and P. Parrilo, · 2015
Later among the works it cites.
“Convergence rate of incremental gradient and newton methods,”
M. Gürbüzbalaban, A. Ozdaglar, and P. Parrilo, · 2015
Later among the works it cites.
“On the convergence rate of incremental aggregated gradient algorithms,”
M. Gurbuzbalaban, A. Ozdaglar, and P. Parrilo, · 2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Accelerating stochastic gradient descent using predictive variance reduction,”
R. Johnson and T. Zhang, · 2013
Cited alongside, same era.
Introductory Lectures on Convex Optimization: A basic course
Y. Nesterov, · 2013
Cited alongside, same era.
“Stochastic gradient descent, weighted sampling, and the randomized kaczmarz algorithm,”
D. Needell, R. Ward, and N. Srebro, · 2014
Cited alongside, same era.
“A note on the non-commutative arithmetic-geometric mean inequality,”
T. Zhang, · 2014
Cited alongside, same era.
“SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives,”
A. Defazio, F. Bach, and S. Lacoste-Julien, · 2014
Cited alongside, same era.
“Finito: A faster, permutable incremental gradient method for big data problems,”
A. Defazio and J. Domke, · 2014
Cited alongside, same era.
“Stochastic gradient descent with finite samples sizes,”
K. Yuan, B. Ying, S. Vlaski, and A. H. Sayed, · 2016
Later among the works it cites.
“Efficient distributed SGD with variance reduction,”
S. De and T. Goldstein, · 2016
Later among the works it cites.
O. Shamir, · 2016
Later among the works it cites.
“On the performance of random reshuffling in stochastic learning,”
B. Ying, K. Yuan, S. Vlaski, and A. H. Sayed, · 2017
Later among the works it cites.
“Variance-reduced stochastic learning under random reshuffling,”
B. Ying, K. Yuan, and A. H. Sayed, · 2017
Later among the works it cites.
“Speeding up distributed machine learning using codes,”
K. Lee, M. Lam, R. Pedarsani, D. Papailiopoulos, and K. Ramchandran, · 2017
Later among the works it cites.
“Convergence of variance-reduced learning under random reshuffling,”
B. Ying, K. Yuan, and A. H. Sayed, · 2018
Closest in time.