Fetching the paper…
Reading the bibliography…
Machine learning models made up of millions or billions of parameters are trained and served on large multi-GPU systems.
Generative communication in linda
David Gelernter · 1985
Earlier work this paper cites.
Efficient all-to-all communication patterns in hypercube and mesh topologies
David S Scott · 1991
Earlier work this paper cites.
Complete exchange on a circuit switched mesh
Shahid H Bokhari and Harry Berryman · 1992
Earlier work this paper cites.
Global combine on mesh architectures with wormhole routing
Michael Barnett, Rick Littlefield, David G Payne, and Robert van de Geijn · 1993
Earlier work this paper cites.
Programming parallel applications in cilk
Charles Leiserson and Aske Plaat · 1997
Earlier work this paper cites.
Optimization of mpi collectives on clusters of large-scale smp’s
Steve Sistare, Rolf Vandevaart, and Eugene Loh · 1999
Earlier work this paper cites.
The hierarchical factor algorithm for all-to-all communication
Peter Sanders and Jesper Larsson Träff · 2002
Earlier work this paper cites.
Improved mpi all-to-all communication on a giganet smp cluster
Jesper Larsson Träff · 2002
Earlier work this paper cites.
Fast collective operations using shared and remote memory access protocols on clusters
Vinod Tipparaju, Jarek Nieplocha, and Dhabaleswar Panda · 2003
Earlier work this paper cites.
Optimization of collective communication operations in mpich
Rajeev Thakur, Rolf Rabenseifner, and William Gropp · 2005
Earlier work this paper cites.
Collective communication: theory, practice, and experience
Ernie Chan, Marcel Heimlich, Avi Purkayastha, and Robert Van De Geijn · 2007
Earlier work this paper cites.
Performance analysis of mpi collective operations
Jelena Pješivac-Grbović, Thara Angskun, George Bosilca, Graham E Fagg, Edgar Gabriel, and Jack J Dongarra · 2007
Earlier work this paper cites.
MPI: A message-passing interface standard version 3.0
Jack Dongarra et al · 2013
Earlier work this paper cites.
Using MPI: Portable Parallel Programming with the Message-Passing Interface
William Gropp, Ewing Lusk, and Anthony Skjellum · 2014
Earlier work this paper cites.
Scalable hierarchical aggregation protocol (sharp): a hardware architecture for efficient data reduction
Richard L Graham, Devendar Bureddy, Pak Lui, Hal Rosenstock, Gilad Shainer, Gil Bloch, Dror Goldenerg, Mike Dubman, Sasha Kotchubievsky, Vladimir Koushnir, et al · 2016
Earlier work this paper cites.
Deep residual learning for image recognition
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2016
Cited alongside, same era.
https://developer.nvidia.com/blog/cooperative-groups/
Cooperative groups: Flexible cuda thread programming, 2017 · 2017
Cited alongside, same era.
Poseidon: An efficient communication architecture for distributed deep learning on GPU clusters
Hao Zhang, Zeyu Zheng, Shizhen Xu, Wei Dai, Qirong Ho, Xiaodan Liang, Zhiting Hu, Jinliang Wei, Pengtao Xie, and Eric P Xing · 2017
Cited alongside, same era.
Horovod: fast and easy distributed deep learning in tensorflow, 2018
Alexander Sergeev and Mike Del Balso · 2018
Cited alongside, same era.
Blueconnect: Novel hierarchical all-reduce on multi-tired network for deep learning
Minsik Cho, Ulrich Finkler, and David Kung · 2019
Cited alongside, same era.
Blueconnect: Decomposing all-reduce for deep learning on heterogeneous network hierarchy
https://www.microsoft.com/en-us/research/blog/deepspeed-accelerating-large-scale-model-inference-and-training-via-system-optimizations-and-compression/
DeepSpeed: Accelerating large-scale model inference and training via system optimizations and compression, 2021 · 2021
Later among the works it cites.
Deeplight: Deep lightweight feature interactions for accelerating ctr predictions in ad serving
Wei Deng, Junwei Pan, Tian Zhou, Deguang Kong, Aaron Flores, and Guang Lin · 2021
Later among the works it cites.
Bluesmpi: Efficient mpi non-blocking alltoall offloading designs on modern bluefield smart nics
Jahanzeb Maqbool Hashmi and Dhabaleswar K Panda · 2021
Later among the works it cites.
Accl: Fpga-accelerated collectives over 100 gbps tcp-ip
Zhenhao He, Daniele Parravicini, Lucian Petrica, Kenneth O’Brien, Gustavo Alonso, and Michaela Blott · 2021
Later among the works it cites.
https://icnc.github.io/
Intel concurrent collections for c++, 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Minsik Cho, Ulrich Finkler, Mauricio Serrano, David Kung, and Hillery Hunter · 2019
Cited alongside, same era.
Tictac: Accelerating distributed deep learning with communication scheduling
Sayed Hadi Hashemi, Sangeetha Abdu Jyothi, and Roy H Campbell · 2019
Cited alongside, same era.
Priority-based parameter propagation for distributed dnn training
Anand Jayarajan, Jinliang Wei, Garth Gibson, Alexandra Fedorova, and Gennady Pekhimenko · 2019
Cited alongside, same era.
A generic communication scheduler for distributed dnn training acceleration
Yanghua Peng, Yibo Zhu, Yangrui Chen, Yixin Bao, Bairen Yi, Chang Lan, Chuan Wu, and Chuanxiong Guo · 2019
Cited alongside, same era.
Scaling distributed machine learning with in-network aggregation
Amedeo Sapio, Marco Canini, Chen-Yu Ho, Jacob Nelson, Panos Kalnis, Changhoon Kim, Arvind Krishnamurthy, Masoud Moshref, Dan RK Ports, and Peter Richtárik · 2019
Cited alongside, same era.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Cited alongside, same era.
Plink: Discovering and exploiting locality for accelerated distributed training on the public cloud
Liang Luo, Peter West, Jacob Nelson, Arvind Krishnamurthy, and Luis Ceze · 2020
Cited alongside, same era.
{ \{ ATP } \} : In-network aggregation for multi-tenant learning
ChonLam Lao, Yanfang Le, Kshiteej Mahajan, Yixi Chen, Wenfei Wu, Aditya Akella, and Michael Swift · 2021
Later among the works it cites.
https://www.nvidia.com/en-us/on-demand/session/gtcspring21-s31578/
Megatron gpt-3 large model inference with triton and onnx runtime, 2021 · 2021
Later among the works it cites.
Scaling distributed machine learning with In-Network aggregation
Amedeo Sapio, Marco Canini, Chen-Yu Ho, Jacob Nelson, Panos Kalnis, Changhoon Kim, Arvind Krishnamurthy, Masoud Moshref, Dan Ports, and Peter Richtarik · 2021
Later among the works it cites.
Ningning Xie, Tamara Norman, Dominik Grewe, and Dimitrios Vytiniotis · 2021
Later among the works it cites.
https://openai.com/blog/ai-and-compute/
AI and compute, 2022 · 2022
Closest in time.
https://developer.nvidia.com/gpudirect
Nvidia gpudirect: Enhancing data movement and access for gpus, 2022 · 2022
Closest in time.
https://github.com/nvidia/nccl
Nvidia collective communication library (nccl), 2022 · 2022
Closest in time.
https://www.alignmentforum.org/posts/GzoWcYibWYwJva8aL/parameter-counts-in-machine-learning
Parameter counts in machine learning, 2022 · 2022
Closest in time.
https://github.com/ROCmSoftwarePlatform/rccl
Rocm communication collectives library (rccl), 2022 · 2022
Closest in time.