Understand
As deep neural networks become more complex and input datasets grow larger, it can take days or even weeks to train a deep neural network to the desired accuracy.
- Therefore, distributed Deep Learning at a massive scale is a critical capability, since it offers the potential to reduce the training time from weeks to hours.
- In this paper, we present a software-hardware co-optimized distributed Deep Learning system that can achieve near-linear scaling up to hundreds of GPUs.
- The core algorithm is a multi-ring communication pattern that provides a good tradeoff between latency and bandwidth and adapts to a variety of system configurations.
Built on
M. Barnett, R. Littlefield, D. Payne, and R. van de Geijn. Global combine on mesh architecture with wormhole routing
1993
Earlier work this paper cites.
DW Walker, JJ Dongarra. MPI: a standard message passing interface
1996
Earlier work this paper cites.
Similar
J. Dean, G. Corrado, R. Monga, K. Chen, M. Devin, M. Mao, M. Ranzato, A. Senior, P. Tucker, K. Yang, Q. Le, A. Ng. Large Scale Distributed Deep Networks
2012
Cited alongside, same era.
Priya Goyal, Piotr Dollar, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, Kaiming He. Accurate, Large Minibatch SGD: Training Imagenet in 1 Hour
Cited in the paper.
Kaiming He, Xiangyu Zhang, Shaoqing Ren, Jian Sun. Deep Residual Learning for Image Recognition
Cited in the paper.
IBM Spectrum MPI: User’s Guide, GC27-8265-01
Cited in the paper.
https://github.com/torch/torch7
Cited in the paper.
Cited in the paper.
Mark Harris. NVIDIA DGX-1: The Fastest Deep Learning System. https://devblogs.nvidia.com/parallelforall/dgx-1-fastest-deep-learning-system
Cited in the paper.
Trishul Chilimbi, Yutaka Suzue, Johnson Apacible, Karthik Kalyanaraman. Project Adam: building an efficient and scalable Deep Learning training system
Cited in the paper.
https://github.com/PPC64/torch-distro
Cited in the paper.
https://github.com/soumith/imagenet-multiGPU.torch
Cited in the paper.
K. He, X. Zhang, S. Ren, and J. Sun. Deep residual learning for image recognition
Cited in the paper.
Priya Goyal, Piotr Dollar, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, Kaiming He. Accurate, Large Minibatch SGD: Training Imagenet in 1 Hour. https://research.fb.com/publications/imagenet1kin1h
Cited in the paper.
Then
Martin Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, Xiaoqiang Zheng. TensorFlow: A system for large-scale machine learning
2015
Later among the works it cites.
Beyond the bibliography
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…