Fetching the paper…
Reading the bibliography…
The performance and efficiency of distributed machine learning (ML) depends significantly on how long it takes for nodes to exchange state changes.
Results of a prototype television bandwidth compression scheme
A. H. Robinson and C. Cherry · 1967
Earlier work this paper cites.
A bridging model for parallel computation
L. Valiant · 1990
Earlier work this paper cites.
Learning multiple layers of features from tiny images
A. Krizhevsky · 2009
Earlier work this paper cites.
Hogwild!: A lock-free approach to parallelizing stochastic gradient descent
B. Recht, C. Re, S. Wright, and F. Niu · 2011
Earlier work this paper cites.
Practical bayesian optimization of machine learning algorithms
J. Snoek, H. Larochelle, and R. P. Adams · 2012
Earlier work this paper cites.
More effective distributed ML via a stale synchronous parallel parameter server
Q. Ho, J. Cipar, H. Cui, S. Lee, J. K. Kim, P. B. Gibbons, G. A. Gibson, G. Ganger, and E. P. Xing · 2013
Earlier work this paper cites.
Project Adam: Building an efficient and scalable deep learning training system
T. Chilimbi, Y. Suzue, J. Apacible, and K. Kalyanaraman · 2014
Earlier work this paper cites.
Scaling distributed machine learning with the parameter server
M. Li, D. G. Andersen, J. W. Park, A. J. Smola, A. Ahmed, V. Josifovski, J. Long, E. J. Shekita, and B.-Y. Su · 2014
Earlier work this paper cites.
1-bit stochastic gradient descent and application to data-parallel distributed training of speech DNNs
F. Seide, H. Fu, J. Droppo, G. Li, and D. Yu · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
K. Simonyan and A. Zisserman · 2014
Earlier work this paper cites.
Distributed machine learning via sufficient factor broadcasting
P. Xie, J. K. Kim, Y. Zhou, Q. Ho, A. Kumar, Y. Yu, and E. Xing · 2014
Earlier work this paper cites.
Deep residual learning for image recognition
K. He, X. Zhang, S. Ren, and J. Sun · 2015
Earlier work this paper cites.
Batch normalization: Accelerating deep network training by reducing internal covariate shift
S. Ioffe and C. Szegedy · 2015
Earlier work this paper cites.
Reducing communication overhead in distributed learning by an order of magnitude (almost)
A. Øland and B. Raj · 2015
Earlier work this paper cites.
Global analytics in the face of bandwidth and regulatory constraints
A. Vulimiri, C. Curino, P. B. Godfrey, T. Jungblut, J. Padhye, and G. Varghese · 2015
Cited alongside, same era.
Managed communication and consistency for fast data-parallel iterative analytics
J. Wei, W. Dai, A. Qiao, Q. Ho, H. Cui, G. R. Ganger, P. B. Gibbons, G. A. Gibson, and E. P. Xing · 2015
Cited alongside, same era.
TensorFlow: A system for large-scale machine learning
M. Abadi, P. Barham, J. Chen, Z. Chen, A. Davis, J. Dean, M. Devin, S. Ghemawat, G. Irving, M. Isard, M. Kudlur, J. Levenberg, R. Monga, S. Moore, D. G. Murray, B. Steiner, P. Tucker, V. Vasudevan, P. Warden, M. Wicke, Y. Yu, and X. Zheng · 2016
Cited alongside, same era.
Towards geo-distributed machine learning
I. Cano, M. Weimer, D. Mahajan, C. Curino, and G. M. Fumarola · 2016
Cited alongside, same era.
Revisiting distributed synchronous SGD
J. Chen, R. Monga, S. Bengio, and R. Jozefowicz · 2016
Cited alongside, same era.
Accurate, large minibatch SGD: training ImageNet in 1 hour
P. Goyal, P. Dollár, R. B. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Later among the works it cites.
Gaia: Geo-distributed machine learning approaching LAN speeds
K. Hsieh, A. Harlap, N. Vijaykumar, D. Konomis, G. R. Ganger, P. B. Gibbons, and O. Mutlu · 2017
Later among the works it cites.
https://software.intel.com/en-us/articles/intel-sdm , 2017
Intel 64 and IA-32 architectures developer’s manual · 2017
Later among the works it cites.
In-datacenter performance analysis of a tensor processing unit
N. P. Jouppi, C. Young, N. Patil, D. Patterson, G. Agrawal, R. Bajwa, et al · 2017
Later among the works it cites.
Overview of China’s Cybersecurity Law
KPMG · 2017
Later among the works it cites.
Deep Gradient Compression: Reducing the communication bandwidth for distributed training
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
GeePS: Scalable deep learning on distributed GPUs with a GPU-specialized parameter server
H. Cui, H. Zhang, G. R. Ganger, P. B. Gibbons, and E. P. Xing · 2016
Cited alongside, same era.
EU Commission and United States agree on new framework for transatlantic data flows: EU-US Privacy Shield
European Commission · 2016
Cited alongside, same era.
Deep Compression: Compressing deep neural networks with pruning, trained quantization and huffman coding
S. Han, H. Mao, and W. J. Dally · 2016
Cited alongside, same era.
Federated learning: Strategies for improving communication efficiency
J. Konečnỳ, H. B. McMahan, F. X. Yu, P. Richtárik, A. T. Suresh, and D. Bacon · 2016
Cited alongside, same era.
Ako: Decentralised deep learning with partial gradient exchange
P. Watcharapichat, V. L. Morales, R. C. Fernandez, and P. Pietzuch · 2016
Cited alongside, same era.
Sparse communication for distributed gradient descent
A. F. Aji and K. Heafield · 2017
Cited alongside, same era.
QSGD: Communication-efficient SGD via gradient quantization and encoding
D. Alistarh, D. Grubic, J. Li, R. Tomioka, and M. Vojnovic · 2017
Cited alongside, same era.
Y. Lin, S. Han, H. Mao, Y. Wang, and W. J. Dally · 2017
Later among the works it cites.
http://tldp.org/HOWTO/Traffic-Control-HOWTO/intro.html , 2017
Linux Traffic Control · 2017
Later among the works it cites.
SGDR: stochastic gradient descent with restarts
I. Loshchilov and F. Hutter · 2017
Later among the works it cites.
Communication-efficient learning of deep networks from decentralized data
B. McMahan, E. Moore, D. Ramage, S. Hampson, and B. A. y Arcas · 2017
Later among the works it cites.
TernGrad: Ternary gradients to reduce communication in distributed deep learning
W. Wen, C. Xu, F. Yan, C. Wu, Y. Wang, Y. Chen, and H. Li · 2017
Later among the works it cites.
Poseidon: An efficient communication architecture for distributed deep learning on GPU clusters
H. Zhang, Z. Zheng, S. Xu, W. Dai, Q. Ho, X. Liang, Z. Hu, J. Wei, P. Xie, and E. P. Xing · 2017
Later among the works it cites.
Learning transferable architectures for scalable image recognition
B. Zoph, V. Vasudevan, J. Shlens, and Q. V. Le · 2017
Later among the works it cites.
http://docs.nvidia.com/cuda/cuda-c-programming-guide/index.html , Jan. 2018
Nvidia cuda c programming guide · 2018
Closest in time.