Fetching the paper…
Reading the bibliography…
Machine learning models are increasingly being trained across multiple GPUs and servers.
Efficient all-to-all communication patterns in hypercube and mesh topologies
David S Scott · 1991
Earlier work this paper cites.
Complete exchange on a circuit switched mesh
Shahid H Bokhari and Harry Berryman · 1992
Earlier work this paper cites.
Global combine on mesh architectures with wormhole routing
Michael Barnett, Rick Littlefield, David G Payne, and Robert van de Geijn · 1993
Earlier work this paper cites.
The communication challenge for mpp: Intel paragon and meiko cs-2
Roger W. Hockney · 1994
Earlier work this paper cites.
Optimization of collective communication operations in mpich
Rajeev Thakur, Rolf Rabenseifner, and William Gropp · 2005
Earlier work this paper cites.
Combinatorial sketching for finite programs
Armando Solar-Lezama, Liviu Tancau, Rastislav Bodik, Sanjit Seshia, and Vijay Saraswat · 2006
Earlier work this paper cites.
Collective communication: theory, practice, and experience
Ernie Chan, Marcel Heimlich, Avi Purkayastha, and Robert Van De Geijn · 2007
Earlier work this paper cites.
Performance analysis of mpi collective operations
Jelena Pješivac-Grbović, Thara Angskun, George Bosilca, Graham E Fagg, Edgar Gabriel, and Jack J Dongarra · 2007
Earlier work this paper cites.
Sketching stencils
Armando Solar-Lezama, Gilad Arnold, Liviu Tancau, Rastislav Bodik, Vijay Saraswat, and Sanjit Seshia · 2007
Earlier work this paper cites.
Program Synthesis by Sketching
Armando Solar-Lezama · 2008
Earlier work this paper cites.
Using program synthesis for social recommendations
Alvin Cheung, Armando Solar-Lezama, and Samuel Madden · 2012
Earlier work this paper cites.
MPI: A message-passing interface standard version 3.0
Jack Dongarra et al · 2013
Earlier work this paper cites.
Achieving high utilization with software-driven wan
Chi-Yao Hong, Srikanth Kandula, Ratul Mahajan, Ming Zhang, Vijay Gill, Mohan Nanduri, and Roger Wattenhofer · 2013
Earlier work this paper cites.
B4: Experience with a globally-deployed software defined wan
Sushant Jain, Alok Kumar, Subhasree Mandal, Joon Ong, Leon Poutievski, Arjun Singh, Subbaiah Venkata, Jim Wanderer, Junlan Zhou, Min Zhu, Jon Zolla, Urs Hölzle, Stephen Stuart, and Amin Vahdat · 2013
Earlier work this paper cites.
Jsketch: sketching for java
Jinseong Jeon, Xiaokang Qiu, Jeffrey S Foster, and Armando Solar-Lezama · 2015
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
Learning to infer graphics programs from hand-drawn images
Kevin Ellis, Daniel Ritchie, Armando Solar-Lezama, and Josh Tenenbaum · 2018
Earlier work this paper cites.
Parameter hub: A rack-scale parameter server for distributed deep neural network training
Liang Luo, Jacob Nelson, Luis Ceze, Amar Phanishayee, and Arvind Krishnamurthy · 2018
Earlier work this paper cites.
Horovod: fast and easy distributed deep learning in tensorflow, 2018
Alexander Sergeev and Mike Del Balso · 2018
Cited alongside, same era.
Radwan: Rate adaptive wide area network
Rachee Singh, Manya Ghobadi, Klaus-Tycho Foerster, Mark Filer, and Phillipa Gill · 2018
Cited alongside, same era.
Blueconnect: Decomposing all-reduce for deep learning on heterogeneous network hierarchy
Minsik Cho, Ulrich Finkler, Mauricio Serrano, David Kung, and Hillery Hunter · 2019
Cited alongside, same era.
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov · 2019
Cited alongside, same era.
https://developer.nvidia.com/blog/massively-scale-deep-learning-training-nccl-2-4
NCCL Tree Algorithm, 2019 · 2019
Cited alongside, same era.
In-network aggregation for shared machine learning clusters
Nadeen Gebara, Manya Ghobadi, and Paolo Costa · 2021
Closest in time.
ATP: In-network aggregation for multi-tenant learning
ChonLam Lao, Yanfang Le, Kshiteej Mahajan, Yixi Chen, Wenfei Wu, Aditya Akella, and Michael Swift · 2021
Closest in time.
https://github.com/microsoft/sccl
Microsoft SCCL, 2021 · 2021
Closest in time.
https://www.microsoft.com/en-us/research/blog/using-deepspeed-and-megatron-to-train-megatron-turing-nlg-530b-the-worlds-largest-and-most-powerful-generative-language-model/
Using deepspeed and megatron to train megatron-turing nlg 530b, the world’s largest and most powerful generative language model · 2021
Closest in time.
https://github.com/NVIDIA/nccl-tests
NCCL Tests, 2021 · 2021
Closest in time.
https://www.nvidia.com/en-us/data-center/dgx-systems/
Nvidia DGX Systems, 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Cited alongside, same era.
Gshard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen · 2020
Cited alongside, same era.
Plink: Discovering and exploiting locality for accelerated distributed training on the public cloud
Liang Luo, Peter West, Jacob Nelson, Arvind Krishnamurthy, and Luis Ceze · 2020
Cited alongside, same era.
Blink: Fast and generic collectives for distributed ml
Guanhua Wang, Shivaram Venkataraman, Amar Phanishayee, Nikhil Devanur, Jorgen Thelin, and Ion Stoica · 2020
Cited alongside, same era.
Is network the bottleneck of distributed training?
Zhen Zhang, Chaokun Chang, Haibin Lin, Yida Wang, Raman Arora, and Xin Jin · 2020
Cited alongside, same era.
https://developer.nvidia.com/gpudirect
GPUDirect RDMA, 2021 · 2021
Cited alongside, same era.
In 18th USENIX Symposium on Networked Systems Design and Implementation (NSDI 21)
Scaling distributed machine learning with In-Network aggregation · 2021
Cited alongside, same era.
Closest in time.
https://www.nvidia.com/en-us/networking/infiniband-adapters/
Nvidia InfiniBand, 2021 · 2021
Closest in time.
https://github.com/nvidia/nccl
Nvidia NCCL, 2021 · 2021
Closest in time.
https://www.nvidia.com/en-us/data-center/nvlink/
Nvidia NVLink and NVSwitch, 2021 · 2021
Closest in time.
https://images.nvidia.com/content/pdf/nvswitch-technical-overview.pdf
NVIDIA NVSWITCH The World’s Highest-Bandwidth On-Node Switch , 2021 · 2021
Closest in time.
Cost-effective cloud edge traffic engineering with cascara
Rachee Singh, Sharad Agarwal, Matt Calder, and Paramvir Bahl · 2021
Closest in time.
Cost-Effective Capacity Provisioning in Wide Area Networks with Shoofly
Rachee Singh, Nikolaj Bjorner, Sharon Shoham, Yawei Yin, John Arnold, and Jamie Gaudette · 2021
Closest in time.
Ningning Xie, Tamara Norman, Dominik Grewe, and Dimitrios Vytiniotis · 2021
Closest in time.
https://github.com/NVIDIA/Megatron-LM , 2022
Megatron-LM · 2022
Closest in time.
https://github.com/NVIDIA/DeepLearningExamples/tree/master/PyTorch/LanguageModeling/Transformer-XL , 2022
Transformer-XL · 2022
Closest in time.
Gurobi Optimizer Reference Manual, 2022
Gurobi Optimization, LLC · 2022
Closest in time.
Software-hardware co-design for fast and scalable training of deep learning recommendation models
Dheevatsa Mudigere, Yuchen Hao, Jianyu Huang, Zhihao Jia, Andrew Tulloch, Srinivas Sridharan, Xing Liu, Mustafa Ozdal, Jade Nie, Jongsoo Park, et al · 2022
Closest in time.