Fetching the paper…
Reading the bibliography…
We develop and analyze a new algorithm for empirical risk minimization, which is the key paradigm for training supervised machine learning models.
H. Robbins and S. Monro, “A stochastic approximation method,” Annals of Mathematical Statistics
1951
Earlier work this paper cites.
R. S. Varga, “Eigenvalues of circulant matrices,” Pacific J. Math
1954
Earlier work this paper cites.
Kluwer Academic Publishers, 2004
Y. Nesterov, Introductory Lectures on Convex Optimization: A Basic Course (Applied Optimization) · 2004
Earlier work this paper cites.
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro, “Robust stochastic approximation approach to stochastic programming,” SIAM Journal on Optimization
2009
Earlier work this paper cites.
M. W. Schmidt, E. V. D. Berg, M. P. Friedlander, and K. Murphy, “Optimizing costly functions with simple constraints: A limited-memory projected quasi-newton algorithm,” International Conference on Artificial Intelligence and Statistics
2009
Earlier work this paper cites.
M. Schmidt, D. Kim, and S. Sra, “Projected Newton-type methods in machine learning,” in Optimization for Machine Learning
2011
Earlier work this paper cites.
C.-C. Chang and C.-J. Lin, “LIBSVM: a library for support vector machines,” ACM Transactions on Intelligent Systems and Technology (TIST)
2011
Earlier work this paper cites.
S. Shalev-Shwartz and T. Zhang, “Stochastic dual coordinate ascent methods for regularized loss minimization,” Journal of Machine Learning Research
2013
Earlier work this paper cites.
R. Johnson and T. Zhang, “Accelerating stochastic gradient descent using predictive variance reduction,” in Advances in Neural Information Processing Systems
2013
Earlier work this paper cites.
M. Takáč, A. Bijral, P. Richtárik, and N. Srebro, “Mini-batch primal and dual methods for SVMs,” in Proceedings of the 30th International Conference on Machine Learning
2013
Earlier work this paper cites.
S. Shalev-Shwartz and T. Zhang, “Accelerated mini-batch stochastic dual coordinate ascent,” in Advances in Neural Information Processing Systems 26
2013
Earlier work this paper cites.
Cambridge University Press, 2014
S. Shalev-Shwartz and S. Ben-David, Understanding machine learning: from theory to algorithms · 2014
Earlier work this paper cites.
P. Richtárik and M. Takáč, “Iteration complexity of randomized block-coordinate descent methods for minimizing a composite function,” Mathematical Programming
2014
Earlier work this paper cites.
A. Defazio, F. Bach, and S. Lacoste-Julien, “SAGA: A fast incremental gradient method with support for non-strongly convex composite objectives,” in Advances in neural information processing systems
2014
Cited alongside, same era.
O. Fercoq, Z. Qu, P. Richtárik, and M. Takáč, “Fast distributed coordinate descent for minimizing non-strongly convex losses,” IEEE International Workshop on Machine Learning for Signal Processing
2014
Cited alongside, same era.
Q. Lin, Z. Lu, and L. Xiao, “An accelerated proximal coordinate gradient method,” in Advances in Neural Information Processing Systems
2014
Cited alongside, same era.
M. W. Schmidt, R. Babanezhad, M. O. Ahmed, A. Defazio, A. Clifton, and A. Sarkar, “Non-uniform stochastic average gradient method for training conditional random fields,” in Proceedings of the Eighteenth International Conference on Artificial Intelligence and Statistics, AISTATS 2015
2015
Cited alongside, same era.
Z. Qu and P. Richtárik, “Coordinate descent with arbitrary sampling I: algorithms and complexity,” Optimization Methods and Software
2016
Later among the works it cites.
Z. Allen-Zhu, Z. Qu, P. Richtárik, and Y. Yuan, “Even faster accelerated coordinate descent using non-uniform sampling,” in International Conference on Machine Learning
2016
Later among the works it cites.
R. M. Gower, D. Goldfarb, and P. Richtárik, “Stochastic block BFGS: Squeezing more curvature out of data,” in Proceedings of the 33rd International Conference on Machine Learning
2016
Later among the works it cites.
P. Moritz, R. Nishihara, and M. I. Jordan, “A linearly-convergent stochastic L-BFGS algorithm,” in International Conference on Artificial Intelligence and Statistics
2016
Later among the works it cites.
Z. Qu, P. Richtárik, M. Takáč, and O. Fercoq, “SDNA: stochastic dual Newton ascent for empirical risk minimization,” in Proceedings of The 33rd International Conference on Machine Learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2015
Cited alongside, same era.
T. Hofmann, A. Lucchi, S. Lacoste-Julien, and B. McWilliams, “Variance reduced stochastic gradient descent with neighbors,” in Advances in Neural Information Processing Systems
2015
Cited alongside, same era.
Z. Qu, P. Richtárik, and T. Zhang, “Quartz: Randomized dual coordinate ascent with arbitrary sampling,” in Advances in Neural Information Processing Systems 28
2015
Cited alongside, same era.
O. Fercoq and P. Richtárik, “Accelerated, parallel and proximal coordinate descent,” SIAM Journal on Optimization
2015
Cited alongside, same era.
M. A. Erdogdu and A. Montanari, “Convergence rates of subampled Newton methods,” in Advances in Neural Information Processing Systems 28
2015
Cited alongside, same era.
P. Richtárik and M. Takáč, “Parallel coordinate descent methods for big data optimization,” Mathematical Programming
2016
Cited alongside, same era.
J. Konečný, J. Lu, P. Richtárik, and M. Takáč, “Mini-batch semi-stochastic gradient descent in the proximal setting,” IEEE Journal of Selected Topics in Signal Processing
2016
Cited alongside, same era.
2016
Cited alongside, same era.
2016
Later among the works it cites.
M. Schmidt, N. Le Roux, and F. Bach, “Minimizing finite sums with the stochastic average gradient,” Mathematical Programming
2017
Later among the works it cites.
J. Konečný and P. Richtárik, “S2GD: Semi-stochastic gradient descent methods,” Frontiers in Applied Mathematics and Statistics
2017
Later among the works it cites.
Z. Allen-Zhu, “Katyusha: The first direct acceleration of stochastic gradient methods,” in Proceedings of the 49th Annual ACM SIGACT Symposium on Theory of Computing
2017
Later among the works it cites.
L. Lei and M. I. Jordan, “Less than a single pass: Stochastically controlled stochastic gradient,” in PMLR: Proceedings of Machine Learning Research (AISTATS 2017)
2017
Later among the works it cites.
R. Tappenden, M. Takáč, and P. Richtárik, “On the complexity of parallel coordinate descent,” Optimization Methods and Software
2018
Closest in time.
R. M. Gower, N. Le Roux, and F. Bach, “Tracking the gradients using the Hessian: A new look at variance reducing stochastic methods,” Proceedings of the 21st International Conference on Artificial Intelligence and Statistics
2018
Closest in time.
2018
Closest in time.