Fetching the paper…
Reading the bibliography…
Training large machine learning models requires a distributed computing approach, with communication of the model updates being the bottleneck.
C. Ma, V. Smith, M. Jaggi, M.I. Jordan, P. Richtárik, and M. Takáč, Adding vs. averaging in distributed primal-dual optimization , in The 32nd International Conference on Machine Learning . 2015, pp. 1973–1982
1982
Earlier work this paper cites.
O. Fercoq, Z. Qu, P. Richtárik, and M. Takáč, Fast distributed coordinate descent for minimizing non-strongly convex losses , IEEE International Workshop on Machine Learning for Signal Processing (2014)
2014
Earlier work this paper cites.
M. Jaggi, V. Smith, M. Takáč, J. Terhorst, S. Krishnan, T. Hofmann, and M.I. Jordan, Communication-efficient distributed dual coordinate ascent , in Advances in Neural Information Processing Systems 27 . 2014
2014
Earlier work this paper cites.
F. Seide, H. Fu, J. Droppo, G. Li, and D. Yu, 1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns , in Fifteenth Annual Conference of the International Speech Communication Association . 2014
2014
Earlier work this paper cites.
O. Shamir, N. Srebro, and T. Zhang, Communication-Efficient Distributed Optimization using an Approximate Newton-type Method , in Proceedings of the 31st International Conference on Machine Learning, PMLR , Vol. 32. 2014, pp. 1000–1008
2014
Earlier work this paper cites.
2015
Earlier work this paper cites.
Y. Zhang and L. Xiao, DiSCO: Distributed Optimization for Self-Concordant Empirical Loss , in Proceedings of the 32nd International Conference on Machine Learning, PMLR , Vol. 37. 2015, pp. 362–370
2015
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
2016
Earlier work this paper cites.
P. Richtárik and M. Takáč, Distributed coordinate descent method for learning with big data , Journal of Machine Learning Research 17 (2016), pp. 1–25
2016
Cited alongside, same era.
D. Alistarh, D. Grubic, J. Li, R. Tomioka, and M. Vojnovic, QSGD: Communication-efficient SGD via gradient quantization and encoding , in Advances in Neural Information Processing Systems . 2017, pp. 1709–1720
2017
Cited alongside, same era.
S. Chunduri, P. Coffman, S. Parker, and K. Kumaran, Performance analysis of mpi on cray xc40 xeon phi system (2017)
2017
Cited alongside, same era.
2017
Cited alongside, same era.
2019
Closest in time.
2019
Closest in time.
E. Gorbunov, D. Kovalev, D. Makarenko, and P. Richtárik, Linearly converging error compensated sgd , Advances in Neural Information Processing Systems 33 (2020), pp. 20889–20900
2020
Closest in time.
2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2017
Cited alongside, same era.
W. Wen, C. Xu, F. Yan, C. Wu, Y. Wang, Y. Chen, and H. Li, Terngrad: Ternary gradients to reduce communication in distributed deep learning , in Advances in Neural Information Processing Systems . 2017, pp. 1509–1519
2017
Cited alongside, same era.
J. Bernstein, Y.X. Wang, K. Azizzadenesheli, and A. Anandkumar, signSGD: Compressed Optimisation for Non-Convex Problems , in Proceedings of the 35th International Conference on Machine Learning , J. Dy and A. Krause, eds., Proceedings of Machine Learning Research Vol. 80, 10–15 Jul, Stockholmsmässan, Stockholm Sweden. PMLR, 2018, pp. 560–569
2018
Cited alongside, same era.
2018
Cited alongside, same era.
2018
Cited alongside, same era.
V. Smith, S. Forte, C. Ma, M. Takáč, M.I. Jordan, and M. Jaggi, CoCoA: A general framework for communication-efficient distributed optimization , Journal of Machine Learning Research 18 (2018), pp. 1–49
2018
Cited alongside, same era.
2020
Closest in time.
D. Kovalev, A. Koloskova, M. Jaggi, P. Richtarik, and S. Stich, A linearly convergent algorithm for decentralized optimization: Sending less bits for free! , in International Conference on Artificial Intelligence and Statistics . PMLR, 2021, pp. 4087–4095
2021
Closest in time.
C. Philippenko and A. Dieuleveut, Preserved central model for faster bidirectional compression in distributed settings , Advances in Neural Information Processing Systems 34 (2021), pp. 2387–2399
2021
Closest in time.
2021
Closest in time.
2021
Closest in time.