Extra: An exact first-order algorithm for decentralized consensus optimization
W. Shi, Q. Ling, G. Wu, and W. Yin · 2015
Cited alongside, same era.
Experiments on parallel training of deep neural network using model averaging
Original
H. Su and H. Chen · 2015
Cited alongside, same era.
Deep learning with elastic averaging sgd
S. Zhang, A. E. Choromanska, and Y. LeCun · 2015
Cited alongside, same era.
Cnn for handwritten arabic digits recognition based on lenet-5
A. El-Sawy, E.-B. Hazem, and M. Loey · 2016
Cited alongside, same era.
Federated learning: Strategies for improving communication efficiency
Original
J. Konečnỳ, H. B. McMahan, F. X. Yu, P. Richtárik, A. T. Suresh, and D. Bacon · 2016
Cited alongside, same era.
Dsa: Decentralized double stochastic averaging gradient algorithm
A. Mokhtari and A. Ribeiro · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision
C. Szegedy, V. Vanhoucke, S. Ioffe, J. Shlens, and Z. Wojna · 2016
Cited alongside, same era.
Sparse communication for distributed gradient descent
A. F. Aji and K. Heafield · 2017
Cited alongside, same era.
Qsgd: Communication-efficient sgd via gradient quantization and encoding
D. Alistarh, D. Grubic, J. Li, R. Tomioka, and M. Vojnovic · 2017
Cited alongside, same era.
Sarah: A novel method for machine learning problems using stochastic recursive gradient
L. M. Nguyen, J. Liu, K. Scheinberg, and M. Takáč · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
A. Paszke, S. Gross, S. Chintala, G. Chanan, E. Yang, Z. DeVito, Z. Lin, A. Desmaison, L. Antiga, and A. Lerer · 2017
Cited alongside, same era.
Don’t use large mini-batches, use local sgd
Original
T. Lin, S. U. Stich, K. K. Patel, and M. Jaggi
Cited in the paper.