Fetching the paper…
Reading the bibliography…
Communication-efficient SGD algorithms, which allow nodes to perform local updates and periodically synchronize local models, are highly effective in improving the speed and scalability of distributed SGD.
Eigenvalues of nonnegative symmetric matrices
Fiedler, M · 1974
Earlier work this paper cites.
Distributed asynchronous deterministic and stochastic gradient optimization algorithms
Tsitsiklis, J., Bertsekas, D., and Athans, M · 1986
Earlier work this paper cites.
Matrix analysis , chapter 5
Horn, R. A. and Johnson, C. R · 1990
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Krizhevsky, A · 2009
Earlier work this paper cites.
Distributed subgradient methods for multi-agent optimization
Nedic, A. and Ozdaglar, A · 2009
Earlier work this paper cites.
Distributed optimization and statistical learning via the alternating direction method of multipliers
Boyd, S., Parikh, N., Chu, E., Peleato, B., and Eckstein, J · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
Recht, B., Re, C., Wright, S., and Niu, F · 2011
Earlier work this paper cites.
Large scale distributed deep networks
Dean, J., Corrado, G., Monga, R., Chen, K., Devin, M., Mao, M., Senior, A., Tucker, P., Yang, K., Le, Q. V., et al · 2012
Earlier work this paper cites.
Optimal distributed online prediction using mini-batches
Dekel, O., Gilad-Bachrach, R., Shamir, O., and Xiao, L · 2012
Earlier work this paper cites.
Dual averaging for distributed optimization: Convergence analysis and network scaling
Duchi, J. C., Agarwal, A., and Wainwright, M. J · 2012
Earlier work this paper cites.
Communication/computation tradeoffs in consensus-based distributed optimization
Tsianos, K., Lawlor, S., and Rabbat, M. G · 2012
Earlier work this paper cites.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
Ghadimi, S. and Lan, G · 2013
Earlier work this paper cites.
Exploiting bounded staleness to speed up big data analytics
Cui, H., Cipar, J., Ho, Q., Kim, J. K., Lee, S., Kumar, A., Wei, J., Dai, W., Ganger, G. R., Gibbons, P. B., et al · 2014
Earlier work this paper cites.
Scaling distributed machine learning with the parameter server
Li, M., Andersen, D. G., Park, J. W., Smola, A. J., Ahmed, A., Josifovski, V., Long, J., Shekita, E. J., and Su, B.-Y · 2014
Earlier work this paper cites.
Proximal algorithms
Parikh, N. and Boyd, S · 2014
Earlier work this paper cites.
Parallel training of dnns with natural gradient and parameter averaging
Povey, D., Zhang, X., and Khudanpur, S · 2014
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
Simonyan, K. and Zisserman, A · 2014
Cited alongside, same era.
Asynchronous parallel stochastic gradient for nonconvex optimization
Lian, X., Huang, Y., Li, Y., and Liu, J · 2015
Cited alongside, same era.
SparkNet: Training deep networks in spark
Moritz, P., Nishihara, R., Stoica, I., and Jordan, M. I · 2015
Cited alongside, same era.
Experiments on parallel training of deep neural network using model averaging
Su, H. and Chen, H · 2015
Collaborative deep learning in fixed topology networks
Jiang, Z., Balu, A., Hegde, C., and Sarkar, S · 2017
Later among the works it cites.
Can decentralized algorithms outperform centralized algorithms? a case study for decentralized parallel stochastic gradient descent
Lian, X., Zhang, C., Zhang, H., Hsieh, C.-J., Zhang, W., and Liu, J · 2017
Later among the works it cites.
Deep gradient compression: Reducing the communication bandwidth for distributed training
Lin, Y., Han, S., Mao, H., Wang, Y., and Dally, W. J · 2017
Later among the works it cites.
Gradient sparsification for communication-efficient distributed optimization
Wangni, J., Wang, J., Liu, J., and Zhang, T · 2017
Later among the works it cites.
TernGrad: Ternary gradients to reduce communication in distributed deep learning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deep learning with elastic averaging SGD
Zhang, S., Choromanska, A. E., and LeCun, Y · 2015
Cited alongside, same era.
Model accuracy and runtime tradeoff in distributed deep learning: A systematic study
Gupta, S., Zhang, W., and Wang, F · 2016
Cited alongside, same era.
How to scale distributed deep learning?
Jin, P. H., Yuan, Q., Iandola, F., and Keutzer, K · 2016
Cited alongside, same era.
Communication-efficient learning of deep networks from decentralized data
McMahan, H. B., Moore, E., Ramage, D., Hampson, S., et al · 2016
Cited alongside, same era.
Asynchrony begets momentum, with an application to deep learning
Mitliagkas, I., Zhang, C., Hadjis, S., and Ré, C · 2016
Cited alongside, same era.
On the convergence of decentralized gradient descent
Yuan, K., Ling, Q., and Yin, W · 2016
Cited alongside, same era.
On nonconvex decentralized gradient descent
Zeng, J. and Yin, W · 2016
Cited alongside, same era.
Wen, W., Xu, C., Yan, F., Wu, C., Wang, Y., Chen, Y., and Li, H · 2017
Later among the works it cites.
Zhou, F. and Cong, G · 2017
Later among the works it cites.
signsgd: compressed optimisation for non-convex problems
Bernstein, J., Wang, Y.-X., Azizzadenesheli, K., and Anandkumar, A · 2018
Closest in time.
Optimization methods for large-scale machine learning
Bottou, L., Curtis, F. E., and Nocedal, J · 2018
Closest in time.
Slow and stale gradients can win the race: Error-runtime trade-offs in distributed SGD
Dutta, S., Joshi, G., Ghosh, S., Dube, P., and Nagpurkar, P · 2018
Closest in time.
Don’t use large mini-batches, use local SGD
Lin, T., Stich, S. U., and Jaggi, M · 2018
Closest in time.
Decentralized consensus algorithm with delayed and stochastic gradients
Sirb, B. and Ye, X · 2018
Closest in time.
Cocoa: A general framework for communication-efficient distributed optimization
Smith, V., Forte, S., Chenxin, M., Takáč, M., Jordan, M. I., and Jaggi, M · 2018
Closest in time.
Local SGD converges fast and communicates little
Stich, S. U · 2018
Closest in time.
Parallel restarted SGD for non-convex optimization with faster convergence and less communication
Yu, H., Yang, S., and Zhu, S · 2018
Closest in time.