Fetching the paper…
Reading the bibliography…
A long-standing problem in the theory of stochastic gradient descent (SGD) is to prove that its without-replacement version RandomShuffle converges faster than the usual with-replacement version.
Gradient methods for the minimisation of functionals
B. T. Polyak · 1963
Earlier work this paper cites.
An adaptive associative memory principle
T. Kohonen · 1974
Earlier work this paper cites.
Problem complexity and method efficiency in optimization
A. Nemirovskii, D. B. Yudin, and E. R. Dawson · 1983
Earlier work this paper cites.
Incremental gradient algorithms with stepsizes bounded away from zero
M. V. Solodov · 1998
Earlier work this paper cites.
Convergence rate of incremental subgradient algorithms
A. Nedić and D. Bertsekas · 2001
Earlier work this paper cites.
Cubic regularization of newton method and its global performance
Y. Nesterov and B. T. Polyak · 2006
Earlier work this paper cites.
Curiously fast convergence of some stochastic gradient descent algorithms
L. Bottou · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro · 2009
Earlier work this paper cites.
Incremental gradient, subgradient, and proximal methods for convex optimization: A survey
D. P. Bertsekas · 2011
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
E. Moulines and F. R. Bach · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
B. Recht, C. Re, S. Wright, and F. Niu · 2011
Earlier work this paper cites.
Stochastic gradient descent tricks
L. Bottou · 2012
Earlier work this paper cites.
Towards a unified architecture for in-rdbms analytics
X. Feng, A. Kumar, B. Recht, and C. Ré · 2012
Cited alongside, same era.
Making gradient descent optimal for strongly convex stochastic optimization
A. Rakhlin, O. Shamir, K. Sridharan, et al · 2012
Cited alongside, same era.
B. Recht and C. Ré · 2012
Cited alongside, same era.
Optimization for machine learning
S. Sra, S. Nowozin, and S. J. Wright · 2012
Cited alongside, same era.
Accelerating stochastic gradient descent using predictive variance reduction
R. Johnson and T. Zhang · 2013
Cited alongside, same era.
Introductory lectures on convex optimization: A basic course , volume 87
Dimension-free iteration complexity of finite sum optimization problems
Y. Arjevani and O. Shamir · 2016
Later among the works it cites.
Optimization methods for large-scale machine learning
L. Bottou, F. E. Curtis, and J. Nocedal · 2016
Later among the works it cites.
Efficient distributed sgd with variance reduction
S. De and T. Goldstein · 2016
Later among the works it cites.
An arithmetic–geometric mean inequality for products of three matrices
A. Israel, F. Krahmer, and R. Ward · 2016
Later among the works it cites.
Random permutations fix a worst case for cyclic coordinate descent
C.-P. Lee and S. J. Wright · 2016
Later among the works it cites.
Stochastic variance reduction for nonconvex optimization
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Nesterov · 2013
Cited alongside, same era.
Stochastic dual coordinate ascent methods for regularized loss minimization
S. Shalev-Shwartz and T. Zhang · 2013
Cited alongside, same era.
Beyond the regret minimization barrier: optimal algorithms for stochastic strongly-convex optimization
E. Hazan and S. Kale · 2014
Cited alongside, same era.
Understanding machine learning: From theory to algorithms
S. Shalev-Shwartz and S. Ben-David · 2014
Cited alongside, same era.
A note on the non-commutative arithmetic-geometric mean inequality
T. Zhang · 2014
Cited alongside, same era.
J. D. Lee, Q. Lin, T. Ma, and T. Yang · 2015
Cited alongside, same era.
Saga: A fast incremental gradient method with support for non-strongly convex composite objectives
A. Defazio, F. Bach, and S. Lacoste-Julien
Cited in the paper.
S. J. Reddi, A. Hefny, S. Sra, B. Poczos, and A. Smola · 2016
Later among the works it cites.
Without-replacement sampling for stochastic gradient methods
O. Shamir · 2016
Later among the works it cites.
Worst-case complexity of cyclic coordinate descent: o ( n 2 ) o(n^{2}) gap with randomized version
R. Sun and Y. Ye · 2016
Later among the works it cites.
When cyclic coordinate descent outperforms randomized coordinate descent
M. Gürbüzbalaban, A. E. Ozdaglar, P. A. Parrilo, and N. D. Vanli · 2017
Later among the works it cites.
Analyzing random permutations for cyclic coordinate descent
S. J. Wright and C.-P. Lee · 2017
Later among the works it cites.
Stochastic learning under random reshuffling
B. Ying, K. Yuan, S. Vlaski, and A. H. Sayed · 2018
Closest in time.