Fetching the paper…
Reading the bibliography…
Model parameter synchronization across GPUs introduces high overheads for data-parallel training at scale.
Edge-disjoint branchings
Jack Edmonds · 1973
Earlier work this paper cites.
On two minimax theorems in graph
Laszlo Lovasz · 1976
Earlier work this paper cites.
Generalized hypercube and hyperbus structures for a computer network
Laxmi N. Bhuyan and Dharma P. Agrawal · 1984
Earlier work this paper cites.
Efficient all-to-all communication patterns in hypercube and mesh topologies
D. Scott · 1991
Earlier work this paper cites.
Global combine on mesh architectures with wormhole routing
M. Barnett, R. Littlefield, D. Payne, and R. van de Geijn · 1993
Earlier work this paper cites.
On global combine operations
Robert van de Geijn · 1994
Earlier work this paper cites.
Packing algorithms for arborescences (and spanning trees) in capacitated graphs
Harold N Gabow and KS Manu · 1998
Earlier work this paper cites.
MagPIe: MPI’s collective communication operations for clustered wide area systems
T. Kielmann, R. F. H. Hofman, H. E. Bal, A. Plaat, and R. A. F. Bhoedjang · 1999
Earlier work this paper cites.
Exploiting hierarchy in parallel computer networks to optimize collective operation performance
N. Karonis, B. de Supinski, I. Foster, W. Gropp, E. Lusk, and J. Bresnahan · 2000
Earlier work this paper cites.
Automatically tuned collective communications
Sathish S. Vadhiyar, Graham E. Fagg, and Jack Dongarra · 2000
Earlier work this paper cites.
Optimization of collective reduction operations
Rolf Rabenseifner · 2004
Earlier work this paper cites.
Optimization of collective communication operations in mpich
Rajeev Thakur, Rolf Rabenseifner, and William Gropp · 2005
Earlier work this paper cites.
Bandwidth efficient all-to-all broadcast on switched clusters
A. Faraj, Pitch Patarasuk, and Xin Yuan · 2008
Earlier work this paper cites.
Bandwidth optimal all-reduce algorithms for clusters of workstations
Pitch Patarasuk and Xin Yuan · 2009
Earlier work this paper cites.
Butterfly mixing: Accelerating incremental-update algorithms on clusters
Huasha Zhao and John Canny · 2013
Earlier work this paper cites.
http://www.ni.com/white-paper/3767/en/ , 2014
PCI Express: An Overview of the PCI Express Standard · 2014
Cited alongside, same era.
Network performance aware mpi collective communication operations in the cloud
Yifan Gong, Bingsheng He, and Jianlong Zhong · 2015
Cited alongside, same era.
Imagenet large scale visual recognition challenge
Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei · 2015
Cited alongside, same era.
Tensorflow: A system for large-scale machine learning
Martin Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2016
Cited alongside, same era.
Machine-aware atomic broadcast trees for multicoress
Stefan Kaestle, Reto Achermann, Roni Haecki, Moritz Hoffmann, Sabela Ramos, and Timothy Roscoe · 2016
Cited alongside, same era.
Poseidon: An efficient communication architecture for distributed deep learning on GPU clusters
Hao Zhang, Zeyu Zheng, Shizhen Xu, Wei Dai, Qirong Ho, Xiaodan Liang, Zhiting Hu, Jinliang Wei, Pengtao Xie, and Eric P. Xing · 2017
Later among the works it cites.
Message Passing Interface
Blaise Barney · 2018
Later among the works it cites.
https://www.nvidia.com/en-us/data-center/dgx-2/ , 2018
NVIDIA DGX-2 · 2018
Later among the works it cites.
Multi-tenant GPU Clusters for Deep Learning Workloads: Analysis and Implications
Myeongjae Jeon, Shivaram Venkataraman, Amar Phanishayee, Junjie Qian, Wencong Xiao, and Fan Yang · 2018
Later among the works it cites.
Training with large minibatches is bad for your health
Yann LeCun · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Near-linear time approximation schemes for some implicit fractional packing problems
Chandra Chekuri and Kent Quanrud · 2017
Cited alongside, same era.
https://www.nvidia.com/en-us/data-center/dgx-1/ , 2017
NVIDIA DGX-1 · 2017
Cited alongside, same era.
Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour
Priya Goyal, Piotr Dollar, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He · 2017
Cited alongside, same era.
IBM Research achieves record deep learning performance with new software technology
Hillery Hunter · 2017
Cited alongside, same era.
Optimized inter-GPU collective operations with NCCL 2
Sylvain Jeaugey · 2017
Cited alongside, same era.
Bringing HPC Techniques to Deep Learning
Andrew Ng · 2017
Cited alongside, same era.
Accelerating machine learning for computer vision
Pieter Noordhuis · 2017
Cited alongside, same era.
Dominic Masters and Carlo Luschi · 2018
Later among the works it cites.
http://images.nvidia.com/content/pdf/nvswitch-technical-overview.pdf , 2018
NVIDIA NVSWITCH · 2018
Later among the works it cites.
https://bit.ly/2k4PXh9 , 2018
Removing roadblocks on the path to 400G and beyond · 2018
Later among the works it cites.
Horovod: fast and easy distributed deep learning in TensorFlow
Alex Sergeev and Mike Del Balso · 2018
Later among the works it cites.
Cachecloud: Towards speed-of-light datacenter communication
Shelby Thomas, Geoffrey M. Voelker, and George Porter · 2018
Later among the works it cites.
https://bit.ly/2lKgAs7 , 2018
Verizon marks milestone with successful 400G technology trial · 2018
Later among the works it cites.
Gandiva: Introspective cluster scheduling for deep learning
Wencong Xiao, Romil Bhardwaj, Ramachandran Ramjee, Muthian Sivathanu, Nipun Kwatra, Zhenhua Han, Pratyush Patel, Xuan Peng, Hanyu Zhao, Quanlu Zhang, Fan Yang, and Lidong Zhou · 2018
Later among the works it cites.
Blueconnect: Decomposing all-reduce for deep learning on heterogeneous network hierarchy
Minsik Cho, Ulrich Finkler, David Kung, and Hillery Hunter · 2019
Closest in time.
Pipedream: Generalized pipeline parallelism for dnn training
Deepak Narayanan, Aaron Harlap, Amar Phanishayee, Vivek Seshadri, Nikhil Devanur, Greg Granger, Phil Gibbons, and Matei Zaharia · 2019
Closest in time.
https://bit.ly/2lFwFQ4 , 2019
Massively Scale Your Deep Learning Training with NCCL 2.4 · 2019
Closest in time.