Fetching the paper…
Reading the bibliography…
We study a distributed consensus-based stochastic gradient descent (SGD) algorithm and show that the rate of convergence involves the spectral properties of two matrices: the standard spectral gap of a weight matrix from the network topology and a new term depending on the spectral norm of the sample covariance matrix of the data.
K. I. Tsianos, S. Lawlor, and M. G. Rabbat, “Communication/computation tradeoffs in consensus-based distributed optimization,” in Advances in Neural Information Processing Systems 25 , P. Bartlett, F. Pereira, C. Burges, L. Bottou, and K. Weinberger, Eds., 2012, pp. 1952–1960
1960
Earlier work this paper cites.
J. A. Blackard and D. J. Dean, “Comparative accuracies of artificial neural networks and discriminant analysis in predicting forest cover types from cartographic variables,” Computers and Electronics in Agriculture , vol. 24, no. 3, pp. 131–151, December 1999. [Online]. Available: http://dx.doi.org/10.1016/S0168-1699(99)00046-0
1999
Earlier work this paper cites.
S. Boyd, L. Xiao, and P. Diaconis, “Fastest mixing markov chain on a graph,” SIAM Review , vol. 46, no. 4, pp. 667–689, 2004. [Online]. Available: http://dx.doi.org/10.1137/S0036144503423264
2004
Earlier work this paper cites.
A. Nedic and A. Ozdaglar, “Distributed subgradient methods for multi-agent optimization,” IEEE Transactions on Automatic Control , vol. 54, no. 1, pp. 48–61, January 2009. [Online]. Available: http://dx.doi.org/10.1109/TAC.2008.2009515
2008
Earlier work this paper cites.
S. Shalev-Shwartz, O. Shamir, N. Srebro, and K. Sridharan, “Stochastic convex optimization,” in Proceedings Conference on Learning Theory , P. Bartlett, F. Pereira, C. Burges, L. Bottou, and K. Weinberger, Eds., 2009
2009
Earlier work this paper cites.
A. S. Bijral and N. Srebro, “On doubly stochastic graph optimization,” in NIPS Workshop on Analyzing Networks and Learning with Graphs , 2009
2009
Earlier work this paper cites.
S. Shalev-Shwartz, Y. Singer, N. Srebro, and A. Cotter, “Pegasos: Primal Estimated sub-GrAdient SOlver for SVM,” Mathematical Programming, Series B , vol. 127, no. 1, pp. 3–30, October 2011. [Online]. Available: http://dx.doi.org/10.1007/s10107-010-0420-4
2011
Earlier work this paper cites.
J. Duchi, A. Agarwal, and M. Wainwright, “Dual averaging for distributed optimization: Convergence analysis and network scaling,” IEEE Transactions on Automatic Control , vol. 57, no. 3, pp. 592–606, March 2011. [Online]. Available: http://dx.doi.org/10.1109/TAC.2011.2161027
2011
Earlier work this paper cites.
S. S. Ram, A. Nedic, and V. V. Veeravalli, “Distributed stochastic subgradient projection algorithms for convex optimization,” Journal of Optimization Theory and Applications , vol. 147, no. 3, pp. 516–545, December 2011. [Online]. Available: http://dx.doi.org/10.1007/s10957-010-9737-7
2011
Earlier work this paper cites.
J. K. Bradley, A. Kyrola, D. Bickson, and C. Guestrin, “Parallel coordinate descent for l 1 l_{1} -regularized loss minimization,” in Proceedings of the 28th International Conference on Machine Learning , ser. JMLR Workshop and Conference Proceedings, L. Getoor and T. Scheffer, Eds., vol. 28, 2011. [Online]. Available: http://www.select.cs.cmu.edu/publications/paperdir/icml2011-bradley-kyr%ola-bickson-guestrin.pdf
2011
Cited alongside, same era.
A. Agarwal and J. C. Duchi, “Distributed delayed stochastic optimization,” in Advances in Neural Information Processing Systems 24 , J. Shawe-Taylor, R. Zemel, P. Bartlett, F. Pereira, and K. Weinberger, Eds., 2011, pp. 873–881. [Online]. Available: http://books.nips.cc/papers/files/nips24/NIPS2011_0574.pdf
2011
Cited alongside, same era.
A. Cotter, O. Shamir, N. Srebro, and K. Sridharan, “Better mini-batch algorithms via accelerated gradient methods,” in Advances in Neural Information Processing Systems 24 , J. Shawe-Taylor, R. Zemel, P. Bartlett, F. Pereira, and K. Weinberger, Eds., 2011, pp. 1647–1655. [Online]. Available: http://papers.nips.cc/paper/4432-better-mini-batch-algorithms-via-accel%erated-gradient-methods
P. Bianchi, G. Fort, and W. Hachem, “Performance of a distributed stochastic approximation algorithm,” IEEE Transactions on Information Theory , vol. 59, no. 11, pp. 7405–7418, 2013. [Online]. Available: http://dx.doi.org/10.1109/TIT.2013.2275131
2013
Later among the works it cites.
M. Takáč, A. Bijral, P. Richtárik, and N. Srebro, “Mini-batch primal and dual methods for SVMs,” in Proceedings of the 30th International Conference on Machine Learning (ICML) , ser. JMLR Workshop and Conference Proceedings, S. Dasgupta and D. McAllester, Eds., vol. 28, 2013, pp. 1022–1030. [Online]. Available: http://jmlr.org/proceedings/papers/v28/takac13.html
2013
Later among the works it cites.
M. Lichman, “UCI machine learning repository,” 2013. [Online]. Available: http://archive.ics.uci.edu/ml
2013
Later among the works it cites.
S. Shalev-Shwartz and S. Ben-David, Understanding Machine Learning: From Theory to Algorithms . Cambridge, UK: Cambridge, 2014
2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2011
Cited alongside, same era.
C.-C. Chang and C.-J. Lin, “Libsvm: A library for support vector machines,” ACM Trans. Intell. Syst. Technol. , vol. 2, no. 3, pp. 27:1–27:27, May 2011. [Online]. Available: http://dx.doi.org/10.1145/1961189.1961199
2011
Cited alongside, same era.
2012
Cited alongside, same era.
O. Dekel, R. Gilad-Bachrach, O. Shamir, and L. Xiao, “Optimal distributed online prediction using mini-batches,” Journal of Machine Learning Research , vol. 13, pp. 165–202, January 2012. [Online]. Available: http://jmlr.org/papers/v13/dekel12a.html
2012
Cited alongside, same era.
Y. Zhang, J. Duchi, and M. Wainwright, “Communication-efficient algorithms for statistical optimization,” in Advances in Neural Information Processing Systems 25 , P. Bartlett, F. Pereira, C. Burges, L. Bottou, and K. Weinberger, Eds., 2012, pp. 1511–1519. [Online]. Available: http://books.nips.cc/papers/files/nips25/NIPS2012_0716.pdf
2012
Cited alongside, same era.
2012
Cited alongside, same era.
2012
Cited alongside, same era.
J. Liu, S. J. Wright, C. Re, V. Bittorf, and S. Sridhar, “An asynchronous parallel stochastic coordinate descent algorithm,” in Proceedings of the 31st International Conference on Machine Learning , ser. JMLR Workshop and Conference Proceedings, L. Getoor and T. Scheffer, Eds., vol. 32, 2014. [Online]. Available: http://jmlr.org/proceedings/papers/v32/liud14.pdf
2014
Later among the works it cites.
O. Shamir, N. Srebro, and T. Zhang, “Communication-efficient distributed optimization using an approximate newton-type method,” in Proceedings of the 31st International Conference on Machine Learning , ser. JMLR Workshop and Conference Proceedings, E. P. Xing and T. Jebara, Eds., vol. 32, 2014, pp. 1000–1008. [Online]. Available: http://jmlr.org/proceedings/papers/v32/shamir14.html
2014
Later among the works it cites.
W. Shi, Q. Ling, G. Wu, and W. Yin, “EXTRA: An exact first-order algorithm for decentralized consensus optimization,” SIAM Journal on Optimization , vol. 25, no. 2, pp. 944–966, 2015
2015
Later among the works it cites.
M. Schmidt, N. L. Roux, and F. Bach, “Minimizing finite sums with the stochastic average gradient,” HAL, Tech. Rep. hal-00860051, January 2015. [Online]. Available: https://hal.inria.fr/hal-00860051v2
2015
Later among the works it cites.
A. Mokhtari and A. Ribeiro, “DSA: decentralized double stochastic averaging gradient algorithm,” Journal of Machine Learning Research , vol. 17, no. 61, pp. 1–35, 2016. [Online]. Available: http://jmlr.org/papers/v17/15-292.html
2016
Closest in time.