S. S. Ram, A. Nedić, and V. V. Veeravalli, “Asynchronous gossip algorithm for stochastic optimization: Constant stepsize analysis,” in Recent Advances in Optimization and its Applications in Engineering . Springer, 2010, pp. 51–60
2010
Earlier work this paper cites.
M. Zinkevich, M. Weimer, L. Li, and A. J. Smola, “Parallelized stochastic gradient descent,” in Advances in neural information processing systems , 2010, pp. 2595–2603
2010
Earlier work this paper cites.
Y. Zhang, J. C. Duchi, and M. J. Wainwright, “Communication-efficient algorithms for statistical optimization,” Journal of Machine Learning Research , vol. 14, pp. 3321–3363, 2013
2013
Earlier work this paper cites.
P. Bianchi, G. Fort, and W. Hachem, “Performance of a distributed stochastic approximation algorithm,” IEEE Transactions on Information Theory , vol. 59, no. 11, pp. 7405–7418, 2013
2013
Earlier work this paper cites.
M. Li, T. Zhang, Y. Chen, and A. J. Smola, “Efficient mini-batch training for stochastic optimization,” in Proceedings of the 20th ACM SIGKDD international conference on Knowledge discovery and data mining , 2014, pp. 661–670
2014
Earlier work this paper cites.
M. Li, D. G. Andersen, J. W. Park, A. J. Smola, A. Ahmed, V. Josifovski, J. Long, E. J. Shekita, and B.-Y. Su, “Scaling distributed machine learning with the parameter server,” in USENIX Symposium on Operating Systems Design and Implementation (OSDI) , 2014
2014
Earlier work this paper cites.
F. Seide, H. Fu, J. Droppo, G. Li, and D. Yu, “1-bit stochastic gradient descent and its application to data-parallel distributed training of speech DNNs,” in Fifteenth Annual Conference of the International Speech Communication Association , 2014
2014
Earlier work this paper cites.
O. Shamir, N. Srebro, and T. Zhang, “Communication-efficient distributed optimization using an approximate Newton-type method,” in International conference on machine learning (ICML) , 2014
2014
Earlier work this paper cites.
J. Kone v
Original
2015
Earlier work this paper cites.
Y. Zhang, J. Duchi, and M. Wainwright, “Divide and conquer kernel ridge regression: a distributed algorithm with minimax optimal rates,” Journal of Machine Learning Research , vol. 16, pp. 3299–3340, 2015
2015
Earlier work this paper cites.
Y. Zhang and X. Lin, “DiSCO: distributed optimization for self-concordant empirical loss,” in International Conference on Machine Learning (ICML) , 2015
2015
Earlier work this paper cites.
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) , 2016
2016
Earlier work this paper cites.
K. Yuan, Q. Ling, and W. Yin, “On the convergence of decentralized gradient descent,” SIAM Journal on Optimization , vol. 26, no. 3, pp. 1835–1854, 2016
2016
Earlier work this paper cites.
B. Sirb and X. Ye, “Consensus optimization with delayed and stochastic gradients on decentralized networks,” in 2016 IEEE International Conference on Big Data (Big Data) . IEEE, 2016, pp. 76–85
2016
Earlier work this paper cites.
V. Smith, S. Forte, C. Ma, M. Takac, M. I. Jordan, and M. Jaggi, “CoCoA: A general framework for communication-efficient distributed optimization,” arXiv preprint arXiv:1611.02189 , 2016
Original
2016
Earlier work this paper cites.
S. J. Reddi, J. Konecnỳ, P. Richtárik, B. Póczós, and A. Smola, “AIDE: fast and communication efficient distributed optimization,” arXiv preprint arXiv:1608.06879 , 2016
Original
2016
Earlier work this paper cites.
J. Konecnỳ, H. B. McMahan, D. Ramage, and P. Richtárik, “Federated optimization: distributed machine learning for on-device intelligence,” arXiv preprint arXiv:1610.02527 , 2016
Original
2016
Earlier work this paper cites.
O. Press and L. Wolf, “Using the output embedding to improve language models,” arXiv preprint arXiv:1608.05859 , 2016
Original
2016
Earlier work this paper cites.
B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas, “Communication-efficient learning of deep networks from decentralized data,” in Artificial Intelligence and Statistics (AISTATS) , 2017
2017
Earlier work this paper cites.