Fetching the paper…
Reading the bibliography…
Several useful variance-reduced stochastic gradient algorithms, such as SVRG, SAGA, Finito, and SAG, have been proposed to minimize empirical risks with linear convergence properties to the exact minimizer.
Introduction to Optimization
B. T. Polyak, · 1987
Earlier work this paper cites.
“Curiously fast convergence of some stochastic gradient descent algorithms,”
L. Bottou, · 2009
Earlier work this paper cites.
“A stochastic gradient method with an exponential convergence rate for finite training sets,”
N. L. Roux, M. Schmidt, and F. R. Bach, · 2012
Earlier work this paper cites.
“Toward a noncommutative arithmetic-geometric mean inequality: Conjectures, case-studies, and consequences,”
B. Recht and C. Ré, · 2012
Earlier work this paper cites.
“Accelerating stochastic gradient descent using predictive variance reduction,”
R. Johnson and T. Zhang, · 2013
Earlier work this paper cites.
“Stochastic dual coordinate ascent methods for regularized loss,”
S. Shalev-Shwartz and Tong Zhang, · 2013
Earlier work this paper cites.
Introductory Lectures on Convex Optimization: A basic course
Y. Nesterov, · 2013
Earlier work this paper cites.
“SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives,”
A. Defazio, F. Bach, and S. Lacoste-Julien, · 2014
Cited alongside, same era.
“Finito: A faster, permutable incremental gradient method for big data problems.,”
A. Defazio, J. Domke, and T. S. Caetano, · 2014
Cited alongside, same era.
“Adaptation, learning, and optimization over networks,”
A. H. Sayed, · 2014
Cited alongside, same era.
“A proximal stochastic gradient method with progressive variance reduction,”
L. Xiao and T. Zhang, · 2014
Cited alongside, same era.
“Why random reshuffling beats stochastic gradient descent,”
M. Gürbüzbalaban, A. Ozdaglar, and P. Parrilo, · 2015
Cited alongside, same era.
“Stop wasting my gradients: Practical SVRG,”
R. Harikandeh, M. O. Ahmed, A. Virani, M. Schmidt, J. Konecny, and S. Sallinen, · 2015
Later among the works it cites.
“Efficient distributed SGD with variance reduction,”
S. De and T. Goldstein, · 2016
Later among the works it cites.
O. Shamir, · 2016
Later among the works it cites.
“Federated optimization: distributed machine learning for on-device intelligence,”
J. Konecny, H.B. McMahan, D. Ramage, and P. Richtárik, · 2016
Later among the works it cites.
“On the performance of random reshuffling in stochastic learning,”
B. Ying, K. Yuan, S. Vlaski, and A. H. Sayed, · 2017
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
M. Gürbüzbalaban, A. Ozdaglar, and P. Parrilo, · 2015
Cited alongside, same era.
“On variance reduction in stochastic gradient descent and its asynchronous variants,”
S. J. Reddi, A. Hefny, S. Sra, B. Poczos, and A. J. Smola, · 2015
Cited alongside, same era.
K. Yuan, B. Ying, and A. H. Sayed, · 2017
Closest in time.