Fetching the paper…
Reading the bibliography…
Training modern deep learning models requires large amounts of computation, often provided by GPUs.
Open MPI: Goals, concept, and design of a next generation MPI implementation
Edgar Gabriel, Graham E. Fagg, George Bosilca, Thara Angskun, Jack J. Dongarra, Jeffrey M. Squyres, Vishal Sahay, Prabhanjan Kambadur, Brian Barrett, Andrew Lumsdaine, Ralph H. Castain, David J. Daniel, Richard L. Graham, and Timothy S. Woodall · 2004
Earlier work this paper cites.
Bandwidth optimal all-reduce algorithms for clusters of workstations
Pitch Patarasuk and Xin Yuan · 2009
Earlier work this paper cites.
Meet Michelangelo: Uber’s machine learning platform
Jeremy Hermann and Mike Del Balso · 2017
Earlier work this paper cites.
Tensorflow benchmarks
TensorFlow Authors · 2017
Earlier work this paper cites.
Accurate, large minibatch SGD: Training ImageNet in 1 hour, 2017, arXiv:1706.02677
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Earlier work this paper cites.
Data parallelism — Wikipedia, the free encyclopedia
Wikipedia · 2017
Earlier work this paper cites.
Bringing HPC techniques to deep learning
Andrew Gibiansky · 2017
Cited alongside, same era.
Message Passing Interface (MPI) Forum Home Page
MPI Forum · 2017
Cited alongside, same era.
baidu-research/tensorflow-allreduce
Andrew Gibiansky and Joel Hestness · 2017
Cited alongside, same era.
NVIDIA collective communications library (NCCL)
NVIDIA · 2017
Cited alongside, same era.
Distributed TensorFlow
TensorFlow Authors · 2017
Cited alongside, same era.
uber/horovod: Distributed training framework for TensorFlow
Alexander Sergeev et al · 2017
Cited alongside, same era.
The Trace Event Profiling Tool (about:tracing)
Chromium Authors · 2017
Later among the works it cites.
TensorFlow Benchmarks
Alexander Sergeev · 2017
Later among the works it cites.
Remote direct memory access — Wikipedia, the free encyclopedia
Wikipedia · 2017
Later among the works it cites.
InfiniBand — Wikipedia, the free encyclopedia
Wikipedia · 2017
Later among the works it cites.
RDMA over Converged Ethernet — Wikipedia, the free encyclopedia
Wikipedia · 2017
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Martín Abadi, Ashish Agarwal, Paul Barham, Eugene Brevdo, Zhifeng Chen, Craig Citro, Greg S. Corrado, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Ian Goodfellow, Andrew Harp, Geoffrey Irving, Michael Isard, Yangqing Jia, Rafal Jozefowicz, Lukasz Kaiser, Manjunath Kudlur, Josh Levenberg, Dan Mane, Rajat Monga, Sherry Moore, Derek Murray, Chris Olah, Mike Schuster, Jonathon Shlens, Benoit Steiner, Ilya Sutskever, Kunal Talwar, Paul Tucker, Vincent Vanhoucke, Vijay Vasudevan, Fernanda Viegas, Oriol Vinyals, Pete Warden, Martin Wattenberg, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng
Cited in the paper.
Tensorflow: A system for large-scale machine learning, 2016b, arXiv:1605.08695
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng
Cited in the paper.