Fetching the paper…
Reading the bibliography…
Collective communication algorithms are an important component of distributed computation.
Efficient all-to-all communication patterns in hypercube and mesh topologies. In The Sixth Distributed Memory Computing Conference, 1991. Proceedings . IEEE Computer Society, 398–399
David S Scott. 1991 · 1991
Earlier work this paper cites.
Complete exchange on a circuit switched mesh. In 1992 Proceedings Scalable High Performance Computing Conference . IEEE Computer Society, 300–301
Shahid H Bokhari and Harry Berryman. 1992 · 1992
Earlier work this paper cites.
Global combine on mesh architectures with wormhole routing. In [1993] Proceedings Seventh International Parallel Processing Symposium . IEEE, 156–162
Michael Barnett, Rick Littlefield, David G Payne, and Robert van de Geijn. 1993 · 1993
Earlier work this paper cites.
Building a high-performance collective communication library. In Supercomputing’94: Proceedings of the 1994 ACM/IEEE Conference on Supercomputing . IEEE, 107–116
Mike Barnett, Satya Gupta, David G Payne, Lance Shuler, Robert van de Geijn, and Jerrell Watts. 1994 · 1994
Earlier work this paper cites.
The communication challenge for MPP: Intel Paragon and Meiko CS-2
Roger W Hockney. 1994 · 1994
Earlier work this paper cites.
Optimization of MPI collectives on clusters of large-scale SMP’s. In Proceedings of the 1999 ACM/IEEE conference on Supercomputing . 23–es
Steve Sistare, Rolf Vandevaart, and Eugene Loh. 1999 · 1999
Earlier work this paper cites.
The hierarchical factor algorithm for all-to-all communication. In European Conference on Parallel Processing . Springer, 799–803
Peter Sanders and Jesper Larsson Träff. 2002 · 2002
Earlier work this paper cites.
Improved MPI all-to-all communication on a Giganet SMP cluster. In European Parallel Virtual Machine/Message Passing Interface Users’ Group Meeting . Springer, 392–400
Jesper Larsson Träff. 2002 · 2002
Earlier work this paper cites.
Fast collective operations using shared and remote memory access protocols on clusters. In Proceedings International Parallel and Distributed Processing Symposium . IEEE, 10–pp
Vinod Tipparaju, Jarek Nieplocha, and Dhabaleswar Panda. 2003 · 2003
Earlier work this paper cites.
Open MPI: Goals, concept, and design of a next generation MPI implementation. In European Parallel Virtual Machine/Message Passing Interface Users’ Group Meeting . Springer, 97–104
Edgar Gabriel, Graham E Fagg, George Bosilca, Thara Angskun, Jack J Dongarra, Jeffrey M Squyres, Vishal Sahay, Prabhanjan Kambadur, Brian Barrett, Andrew Lumsdaine, et al · 2004
Earlier work this paper cites.
Optimization of collective communication operations in MPICH
Rajeev Thakur, Rolf Rabenseifner, and William Gropp. 2005 · 2005
Cited alongside, same era.
Collective communication: theory, practice, and experience
Ernie Chan, Marcel Heimlich, Avi Purkayastha, and Robert Van De Geijn. 2007 · 2007
Cited alongside, same era.
Performance analysis of MPI collective operations
Jelena Pješivac-Grbović, Thara Angskun, George Bosilca, Graham E Fagg, Edgar Gabriel, and Jack J Dongarra. 2007 · 2007
Cited alongside, same era.
Z3: An Efficient SMT Solver. In TACAS
Leonardo de Moura and Nikolaj Bjørner. 2008 · 2008
Cited alongside, same era.
MPI: A message-passing interface standard version 3.0
Jack Dongarra et al · 2013
Cited alongside, same era.
Poseidon: An efficient communication architecture for distributed deep learning on GPU clusters. In 2017 USENIX Annual Technical Conference (USENIX ATC 17) . 181–193
Evaluating Modern GPU Interconnect: PCIe, NVLink, NV-SLI, NVSwitch and GPUDirect
A. Li, S. L. Song, J. Chen, J. Li, X. Liu, N. R. Tallent, and K. J. Barker. 2020 · 2019
Later among the works it cites.
A generic communication scheduler for distributed dnn training acceleration. In Proceedings of the 27th ACM Symposium on Operating Systems Principles . 16–29
Yanghua Peng, Yibo Zhu, Yangrui Chen, Yixin Bao, Bairen Yi, Chang Lan, Chuan Wu, and Chuanxiong Guo. 2019 · 2019
Later among the works it cites.
AMD Radeon Instinct MI50 Accelerator
AMD Radeon Instinct MI50 2020 · 2020
Closest in time.
ROCm Communication Collectives Library
AMD RCCL Library 2020 · 2020
Closest in time.
Google Cloud TPU
Google TPU 2020 · 2020
Closest in time.
Graphcore Intelligence Processing Unit
Graphcore IPU 2020 · 2020
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Hao Zhang, Zeyu Zheng, Shizhen Xu, Wei Dai, Qirong Ho, Xiaodan Liang, Zhiting Hu, Jinliang Wei, Pengtao Xie, and Eric P Xing. 2017 · 2017
Cited alongside, same era.
Horovod: fast and easy distributed deep learning in TensorFlow
Alexander Sergeev and Mike Del Balso. 2018 · 2018
Cited alongside, same era.
BlueConnect: Decomposing all-reduce for deep learning on heterogeneous network hierarchy
Minsik Cho, Ulrich Finkler, Mauricio Serrano, David Kung, and Hillery Hunter. 2019 · 2019
Cited alongside, same era.
TicTac: Accelerating distributed deep learning with communication scheduling
Sayed Hadi Hashemi, Sangeetha Abdu Jyothi, and Roy H Campbell. 2019 · 2019
Cited alongside, same era.
Priority-based parameter propagation for distributed DNN training
Anand Jayarajan, Jinliang Wei, Garth Gibson, Alexandra Fedorova, and Gennady Pekhimenko. 2019 · 2019
Cited alongside, same era.
PLink: Discovering and Exploiting Locality for Accelerated Distributed Training on the public Cloud
Liang Luo, Peter West, Jacob Nelson, Arvind Krishnamurthy, and Luis Ceze. 2020 · 2020
Closest in time.
NVIDIA Collective Communications Library
NVIDIA NCCL Library 2020 · 2020
Closest in time.
Unified Communication X
UCX 2020 · 2020
Closest in time.
Blink: Fast and Generic Collectives for Distributed ML. In Conference on Machine Learning and Systems (MLSys 2020)
Guanhua Wang, Shivaram Venkataraman, Amar Phanishayee, Jorgen Thelin, Nikhil Devanur, and Ion Stoica. 2020 · 2020
Closest in time.