Fetching the paper…
Reading the bibliography…
Most of today's distributed machine learning systems assume {\em reliable networks}: whenever two machines exchange information (e.g., gradients or models), the network should guarantee the delivery of the message.
Optimization of collective communication operations in mpich
R. Thakur, R. Rabenseifner, and W. Gropp · 2005
Earlier work this paper cites.
Randomized gossip algorithms
S. Boyd, A. Ghosh, B. Prabhakar, and D. Shah · 2006
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky and G. Hinton · 2009
Earlier work this paper cites.
Consensus conditions of multi-agent systems with time-varying topologies and stochastic communication noises
T. Li and J.-F. Zhang · 2010
Earlier work this paper cites.
Distributed delayed stochastic optimization
A. Agarwal and J. C. Duchi · 2011
Earlier work this paper cites.
Distributed subgradient methods for convex optimization over random networks
I. Lobel and A. Ozdaglar · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
B. Recht, C. Re, S. Wright, and F. Niu · 2011
Earlier work this paper cites.
Asynchronous stochastic gradient descent for dnn training
S. Zhang, C. Zhang, Z. You, R. Zheng, and B. Xu · 2013
Earlier work this paper cites.
Scaling distributed machine learning with the parameter server
M. Li, D. G. Andersen, J. W. Park, A. J. Smola, A. Ahmed, V. Josifovski, J. Long, E. J. Shekita, and B.-Y. Su · 2014
Earlier work this paper cites.
Asynchronous parallel stochastic gradient for nonconvex optimization
X. Lian, Y. Huang, Y. Li, and J. Liu · 2015
Earlier work this paper cites.
Distributed optimization over time-varying directed graphs
A. Nedić and A. Olshevsky · 2015
Earlier work this paper cites.
Adadelay: Delay adaptive distributed stochastic convex optimization
S. Sra, A. W. Yu, M. Li, and A. J. Smola · 2015
Earlier work this paper cites.
Deep learning with elastic averaging sgd
S. Zhang, A. E. Choromanska, and Y. LeCun · 2015
Earlier work this paper cites.
Tensorflow: A system for large-scale machine learning
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, et al · 2016
Earlier work this paper cites.
Gossip training for deep learning
M. Blot, D. Picard, M. Cord, and N. Thome · 2016
Earlier work this paper cites.
Gossip dual averaging for decentralized optimization of pairwise functions
I. Colin, A. Bellet, J. Salmon, and S. Clémençon · 2016
Cited alongside, same era.
Unwrapping admm: efficient distributed computing via transpose reduction
T. Goldstein, G. Taylor, K. Barabin, and K. Sayre · 2016
Cited alongside, same era.
Nestt: A nonconvex primal-dual splitting method for distributed and stochastic optimization
D. Hajinezhad, M. Hong, T. Zhao, and Z. Wang · 2016
Cited alongside, same era.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2016
Cited alongside, same era.
How to scale distributed deep learning?
P. H. Jin, Q. Yuan, F. Iandola, and K. Keutzer · 2016
Cited alongside, same era.
Optimal algorithms for smooth and strongly convex distributed optimization in networks
K. Scaman, F. Bach, S. Bubeck, Y. T. Lee, and L. Massoulié · 2017
Later among the works it cites.
Terngrad: Ternary gradients to reduce communication in distributed deep learning
W. Wen, C. Xu, F. Yan, C. Wu, Y. Wang, Y. Chen, and H. Li · 2017
Later among the works it cites.
Adaptive consensus admm for distributed optimization
Z. Xu, G. Taylor, H. Li, M. Figueiredo, X. Yuan, and T. Goldstein · 2017
Later among the works it cites.
arXiv preprint arXiv:1806.08054 , 2018
Error compensated quantized sgd and its applications to large-scale distributed optimization · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
R. Leblond, F. Pedregosa, and S. Lacoste-Julien · 2016
Cited alongside, same era.
Communication-efficient learning of deep networks from decentralized data
H. B. McMahan, E. Moore, D. Ramage, S. Hampson, et al · 2016
Cited alongside, same era.
Cntk: Microsoft’s open-source deep-learning toolkit
F. Seide and A. Agarwal · 2016
Cited alongside, same era.
Consensus optimization with delayed and stochastic gradients on decentralized networks
B. Sirb and X. Ye · 2016
Cited alongside, same era.
Efficient distributed learning with sparsity
J. Wang, M. Kolar, N. Srebro, and T. Zhang · 2016
Cited alongside, same era.
Machine learning with adversaries: Byzantine tolerant gradient descent
P. Blanchard, R. Guerraoui, J. Stainer, et al · 2017
Cited alongside, same era.
Accurate, large minibatch sgd: training imagenet in 1 hour
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Cited alongside, same era.
D. Alistarh, Z. Allen-Zhu, and J. Li · 2018
Closest in time.
Gossipgrad: Scalable deep learning using gossip communication based asynchronous gradient descent
J. Daily, A. Vishnu, C. Siegel, T. Warfel, and V. Amatya · 2018
Closest in time.
Training dnns with hybrid block floating point
M. Drumond, T. LIN, M. Jaggi, and B. Falsafi · 2018
Closest in time.
Cola: Decentralized linear learning
L. He, A. Bian, and M. Jaggi · 2018
Closest in time.
Don’t use large mini-batches, use local sgd
T. Lin, S. U. Stich, and M. Jaggi · 2018
Closest in time.
Sparcml: High-performance sparse communication for machine learning
C. Renggli, D. Alistarh, and T. Hoefler · 2018
Closest in time.
Local sgd converges fast and communicates little
S. U. Stich · 2018
Closest in time.
Sparsified sgd with memory
S. U. Stich, J.-B. Cordonnier, and M. Jaggi · 2018
Closest in time.
Byzantine-robust distributed learning: Towards optimal statistical rates
D. Yin, Y. Chen, K. Ramchandran, and P. Bartlett · 2018
Closest in time.
Distributed asynchronous optimization with unbounded delays: How slow can you go?
Z. Zhou, P. Mertikopoulos, N. Bambos, P. Glynn, Y. Ye, L.-J. Li, and L. Fei-Fei · 2018
Closest in time.