Fetching the paper…
Reading the bibliography…
Most commonly used distributed machine learning systems are either synchronous or centralized asynchronous.
Gossip algorithms: Design, analysis and applications
S. Boyd, A. Ghosh, B. Prabhakar, and D. Shah · 2005
Earlier work this paper cites.
Consensus and cooperation in networked multi-agent systems
R. Olfati-Saber, J. A. Fax, and R. M. Murray · 2007
Earlier work this paper cites.
A distributed consensus protocol for clock synchronization in wireless sensor network
L. Schenato and G. Gamba · 2007
Earlier work this paper cites.
Randomized consensus algorithms over large scale networks
F. Fagnani and S. Zampieri · 2008
Earlier work this paper cites.
Broadcast gossip algorithms for consensus
T. C. Aysal, M. E. Yildiz, A. D. Sarwate, and A. Scaglione · 2009
Earlier work this paper cites.
Distributed subgradient methods for multi-agent optimization
A. Nedic and A. Ozdaglar · 2009
Earlier work this paper cites.
Robust stochastic approximation approach to stochastic programming
A. Nemirovski, A. Juditsky, G. Lan, and A. Shapiro · 2009
Earlier work this paper cites.
Bandwidth optimal all-reduce algorithms for clusters of workstations
P. Patarasuk and X. Yuan · 2009
Earlier work this paper cites.
Distributed subgradient projection algorithm for convex optimization
S. S. Ram, A. Nedic, and V. V. Veeravalli · 2009
Earlier work this paper cites.
Gossip consensus algorithms via quantized communication
R. Carli, F. Fagnani, P. Frasca, and S. Zampieri · 2010
Earlier work this paper cites.
A gossip algorithm for convex consensus optimization over networks
J. Lu, C. Y. Tang, P. R. Regier, and T. D. Bow · 2010
Earlier work this paper cites.
Asynchronous gossip algorithm for stochastic optimization: Constant stepsize analysis
S. S. Ram, A. Nedić, and V. V. Veeravalli · 2010
Earlier work this paper cites.
Distributed stochastic subgradient projection algorithms for convex optimization
S. Sundhar Ram, A. Nedić, and V. Veeravalli · 2010
Earlier work this paper cites.
Distributed delayed stochastic optimization
A. Agarwal and J. C. Duchi · 2011
Earlier work this paper cites.
Non-asymptotic analysis of stochastic approximation algorithms for machine learning
E. Moulines and F. R. Bach · 2011
Earlier work this paper cites.
Hogwild: A lock-free approach to parallelizing stochastic gradient descent
B. Recht, C. Re, S. Wright, and F. Niu · 2011
Earlier work this paper cites.
Distributed asynchronous constrained stochastic optimization
K. Srivastava and A. Nedic · 2011
Earlier work this paper cites.
Optimal distributed online prediction using mini-batches
O. Dekel, R. Gilad-Bachrach, O. Shamir, and L. Xiao · 2012
Cited alongside, same era.
Performance of a distributed stochastic approximation algorithm
P. Bianchi, G. Fort, and W. Hachem · 2013
Cited alongside, same era.
Stochastic first-and zeroth-order methods for nonconvex stochastic programming
S. Ghadimi and G. Lan · 2013
Cited alongside, same era.
Gpu asynchronous stochastic gradient descent to speed up neural network training
T. Paine, H. Jin, J. Yang, Z. Lin, and T. Huang · 2013
Cited alongside, same era.
Asynchronous distributed admm for consensus optimization
R. Zhang and J. Kwok · 2014
Cited alongside, same era.
Fast multi-gpu collectives with nccl, 2016
N. Luehr · 2016
Later among the works it cites.
DSA: decentralized double stochastic averaging gradient algorithm
A. Mokhtari and A. Ribeiro · 2016
Later among the works it cites.
CNTK: Microsoft’s open-source deep-learning toolkit
F. Seide and A. Agarwal · 2016
Later among the works it cites.
Consensus optimization with delayed and stochastic gradients on decentralized networks
B. Sirb and X. Ye · 2016
Later among the works it cites.
Efficient distributed online prediction and stochastic optimization with approximate distributed averaging
K. I. Tsianos and M. G. Rabbat · 2016
Later among the works it cites.
Decentralized rls with data-adaptive censoring for regressions over large-scale networks
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
N. S. Aybat, Z. Wang, T. Lin, and S. Ma · 2015
Cited alongside, same era.
Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems
T. Chen, M. Li, Y. Li, M. Lin, N. Wang, M. Wang, T. Xiao, B. Xu, C. Zhang, and Z. Zhang · 2015
Cited alongside, same era.
Asynchronous parallel stochastic gradient for nonconvex optimization
X. Lian, Y. Huang, Y. Li, and J. Liu · 2015
Cited alongside, same era.
MPI AllReduce, 2015
MPI contributors · 2015
Cited alongside, same era.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2015
Cited alongside, same era.
CIFAR VGG in Torch
S. Zagoruyko · 2015
Cited alongside, same era.
Deep learning with elastic averaging SGD
S. Zhang, A. E. Choromanska, and Y. LeCun · 2015
Cited alongside, same era.
Z. Wang, Z. Yu, Q. Ling, D. Berberidis, and G. B. Giannakis · 2016
Later among the works it cites.
Decentralized consensus optimization with asynchrony and delays
T. Wu, K. Yuan, Q. Ling, W. Yin, and A. H. Sayed · 2016
Later among the works it cites.
On the convergence of decentralized gradient descent
K. Yuan, Q. Ling, and W. Yin · 2016
Later among the works it cites.
Model accuracy and runtime tradeoff in distributed deep learning: A systematic study
W. Zhang, S. Gupta, and F. Wang · 2016
Later among the works it cites.
ResNet in Torch
FAIR · 2017
Closest in time.
Accurate, large minibatch SGD: training imagenet in 1 hour
P. Goyal, P. Dollár, R. B. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Closest in time.
Communication-efficient algorithms for decentralized and stochastic optimization
G. Lan, S. Lee, and Y. Zhou · 2017
Closest in time.
Z. Li, W. Shi, and M. Yan · 2017
Closest in time.
Can decentralized algorithms outperform centralized algorithms? A case study for decentralized parallel stochastic gradient descent, 2017
X. Lian, C. Zhang, H. Zhang, C.-J. Hsieh, W. Zhang, and J. Liu · 2017
Closest in time.
Wildfire: Approximate synchronization of parameters in distributed deep learning
R. Nair and S. Gupta · 2017
Closest in time.
Gadei: On scale-up training as a service for deep learning
W. Zhang, M. Feng, Y. Zheng, Y. Ren, Y. Wang, J. Liu, P. Liu, B. Xiang, L. Zhang, B. Zhou, and F. Wang · 2017
Closest in time.
D2: Decentralized training over decentralized data
H. Tang, X. Lian, M. Yan, C. Zhang, and J. Liu · 2018
Closest in time.