Fetching the paper…
Reading the bibliography…
Mini-batch optimization has proven to be a powerful paradigm for large-scale learning.
H. Robbins and S. Monro, “A stochastic approximation method,” Annals of Mathematical Statistics , pp. 400–407, 1951
1951
Earlier work this paper cites.
R. Rockafellar and R. Wets, “On the interchange of subdifferentiation and conditional expectation for convex functionals,” Stochastics: An International Journal of Probability and Stochastic Processes , vol. 7, no. 3, pp. 173–182, 1982
1982
Earlier work this paper cites.
J. N. Tsitsiklis, D. P. Bertsekas, and M. Athans, “Distributed asynchronous deterministic and stochastic gradient optimization algorithms,” IEEE Transactions on Automatic Control , vol. 31, no. 9, pp. 803–812, 1986
1986
Earlier work this paper cites.
K. P. Bennett and O. L. Mangasarian, “Robust linear programming discrimination of two linearly inseparable sets,” Optimization Methods and Software , vol. 1, no. 1, pp. 23–34, 1992
1992
Earlier work this paper cites.
G. Chen and M. Teboulle, “Convergence analysis of a proximal-like minimization algorithm using Bregman functions,” SIAM Journal on Optimization , vol. 3, no. 3, pp. 538–543, 1993
1993
Earlier work this paper cites.
R. Tibshirani, “Regression shrinkage and selection via the Lasso,” Journal of the Royal Statistical Society: Series B (Methodological) , pp. 267–288, 1996
1996
Earlier work this paper cites.
A. Nedić, D. P. Bertsekas, and V. S. Borkar, “Distributed asynchronous incremental subgradient methods,” Studies in Computational Mathematics , vol. 8, pp. 381–407, 2001
2001
Earlier work this paper cites.
A. Beck and M. Teboulle, “Mirror descent and nonlinear projected subgradient methods for convex optimization,” Operations Research Letters , vol. 31, no. 3, pp. 167–175, 2003
2003
Earlier work this paper cites.
D. D. Lewis, Y. Yang, T. G. Rose, and F. Li, “RCV1: A new benchmark collection for text categorization research,” Journal of Machine Learning Research , vol. 5, pp. 361–397, 2004
2004
Earlier work this paper cites.
H. Zou and T. Hastie, “Regularization and variable selection via the elastic net,” Journal of the Royal Statistical Society: Series B (Statistical Methodology) , vol. 67, no. 2, pp. 301–320, 2005
2005
Earlier work this paper cites.
T. Hastie, R. Tibshirani, and J. Friedman, The elements of statistical learning . Springer, 2009
2009
Earlier work this paper cites.
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro, “Robust stochastic approximation approach to stochastic programming,” SIAM Journal on Optimization , vol. 19, no. 4, pp. 1574–1609, 2009
2009
Earlier work this paper cites.
L. Xiao, “Dual averaging method for regularized stochastic learning and online optimization,” Advances in Neural Information Processing Systems , pp. 2116–2124, 2009
2009
Earlier work this paper cites.
C. Hu, W. Pan, and J. T. Kwok, “Accelerated gradient methods for stochastic optimization and online learning,” Advances in Neural Information Processing Systems , pp. 781–789, 2009
2009
Earlier work this paper cites.
M. Zinkevich, J. Langford, and A. J. Smola, “Slow learners are fast,” Advances in Neural Information Processing Systems , pp. 2331–2339, 2009
2009
Earlier work this paper cites.
M. Zinkevich, M. Weimer, L. Li, and A. J. Smola, “Parallelized stochastic gradient descent,” Advances in Neural Information Processing Systems , pp. 2595–2603, 2010
2010
Cited alongside, same era.
P. Tseng, “Approximation accuracy, gradient methods, and error bound for structured convex optimization,” Mathematical Programming , vol. 125, no. 2, pp. 263–295, 2010
2010
Cited alongside, same era.
J. Duchi, S. Shalev-Shwartz, Y. Singer, and A. Tewari, “Composite objective mirror descent,” In Annual Conference on Learning Theory (COLT) , 2010
2010
Cited alongside, same era.
S. Shalev-Shwartz and A. Tewari, “Stochastic methods for l 1 l_{1} -regularized loss minimization,” Journal of Machine Learning Research , vol. 12, pp. 1865–1892, 2011
2011
Cited alongside, same era.
I. Lobel and A. Ozdaglar, “Distributed subgradient methods for convex optimization over random networks,” IEEE Transactions on Automatic Control , vol. 56, no. 6, pp. 1291–1306, 2011
P. Bianchi and J. Jakubowicz, “Convergence of a multi-agent projected stochastic gradient algorithm for non-convex optimization,” IEEE Transactions on Automatic Control , vol. 58, no. 2, pp. 391–405, 2013
2013
Later among the works it cites.
A. Nedić and S. Lee, “On stochastic subgradient mirror-descent algorithm with weighted averaging,” SIAM Journal on Optimization , vol. 24, no. 1, pp. 84–107, 2014
2014
Later among the works it cites.
D. Needell, R. Ward, and N. Srebro, “Stochastic gradient descent, weighted sampling, and the randomized Kaczmarz algorithm,” Advances in Neural Information Processing Systems , pp. 1017–1025, 2014
2014
Later among the works it cites.
B. McMahan and M. Streeter, “Delay-tolerant algorithms for asynchronous distributed online learning,” Advances in Neural Information Processing Systems , pp. 2915–2923, 2014
2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2011
Cited alongside, same era.
F. Niu, B. Recht, C. Ré, and S. J. Wright, “Hogwild!: A lock-free approach to parallelizing stochastic gradient descent.” Advances in Neural Information Processing Systems , pp. 693–701, 2011
2011
Cited alongside, same era.
H. H. Bauschke and P. L. Combettes, Convex analysis and monotone operator theory in Hilbert spaces . Springer, 2011
2011
Cited alongside, same era.
A. Cotter, O. Shamir, N. Srebro, and K. Sridharan, “Better mini-batch algorithms via accelerated gradient methods,” Advances in Neural Information Processing Systems , pp. 1647–1655, 2011
2011
Cited alongside, same era.
G. Lan, “An optimal method for stochastic composite optimization,” Mathematical Programming , vol. 133, pp. 365–397, 2012
2012
Cited alongside, same era.
S. Ghadimi and G. Lan, “Optimal stochastic approximation algorithms for strongly convex stochastic composite optimization I: A generic algorithmic framework,” SIAM Journal on Optimization , vol. 22, no. 4, pp. 1469–1492, 2012
2012
Cited alongside, same era.
K. Tsianos, S. Lawlor, and M. G. Rabbat, “Communication/computation tradeoffs in consensus-based distributed optimization,” Advances in Neural Information Processing Systems , pp. 1943–1951, 2012
2012
Cited alongside, same era.
O. Dekel, R. Gilad-Bachrach, O. Shamir, and L. Xiao, “Optimal distributed online prediction using mini-batches,” Journal of Machine Learning Research , vol. 13, no. 1, pp. 165–202, 2012
2012
Cited alongside, same era.
2014
Later among the works it cites.
M. Jaggi, V. Smith, M. Takác, J. Terhorst, S. Krishnan, T. Hofmann, and M. I. Jordan, “Communication-efficient distributed dual coordinate ascent,” Advances in Neural Information Processing Systems , pp. 3068–3076, 2014
2014
Later among the works it cites.
2014
Later among the works it cites.
R. Zhang and J. Kwok, “Asynchronous distributed ADMM for consensus optimization,” Proceedings of the 31st International Conference on Machine Learning (ICML) , pp. 1701–1709, 2014
2014
Later among the works it cites.
S. J. Wright, “Coordinate descent algorithms,” arXiv preprint arXiv:1502.04759 , 2014
2014
Later among the works it cites.
J. Liu, S. J. Wright, C. Ré, V. Bittorf, and S. Sridhar, “An asynchronous parallel stochastic coordinate descent algorithm,” Proceedings of the 31st International Conference on Machine Learning (ICML) , pp. 469–477, 2014
2014
Later among the works it cites.
2015
Closest in time.
2015
Closest in time.
P. Richtárik and M. Takác, “Parallel coordinate descent methods for big data optimization,” Mathematical Programming , pp. 1–52, 2015
2015
Closest in time.
J. Liu and S. J. Wright, “Asynchronous stochastic coordinate descent: Parallelism and convergence properties,” SIAM Journal on Optimization , vol. 25, no. 1, pp. 351–376, 2015
2015
Closest in time.