Fetching the paper…
Reading the bibliography…
Zero Redundancy Optimizer (ZeRO) has been used to train a wide range of large language models on massive GPUs clusters due to its ease of use, efficiency, and good scalability.
Reducing BERT Pre-Training Time from 3 Days to 76 Minutes
Yang You, Jing Li, Jonathan Hseu, Xiaodan Song, James Demmel, and Cho-Jui Hsieh. 2019 · 1904
Earlier work this paper cites.
Optimization of collective communication operations in MPICH
Rajeev Thakur, Rolf Rabenseifner, and William Gropp. 2005 · 2005
Earlier work this paper cites.
Collective communication: theory, practice, and experience
Ernie Chan, Marcel Heimlich, Avi Purkayastha, and Robert Van De Geijn. 2007 · 2007
Earlier work this paper cites.
Large scale distributed deep networks
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Marc’aurelio Ranzato, Andrew Senior, Paul Tucker, Ke Yang, et al · 2012
Earlier work this paper cites.
1-bit stochastic gradient descent and its application to data-parallel distributed training of speech dnns. In Fifteenth annual conference of the international speech communication association
Frank Seide, Hao Fu, Jasha Droppo, Gang Li, and Dong Yu. 2014 · 2014
Earlier work this paper cites.
8-bit approximations for parallelism in deep learning
Tim Dettmers. 2015 · 2015
Earlier work this paper cites.
Scalable distributed DNN training using commodity GPU cloud computing
Nikko Ström. 2015 · 2015
Earlier work this paper cites.
Communication Quantization for Data-Parallel Training of Deep Neural Networks. In Proceedings of the Workshop on Machine Learning in High Performance Computing Environments (Salt Lake City, Utah) (MLHPC ’16) . IEEE Press, 1–8
Nikoli Dryden, Tim Moon, Sam Ade Jacobs, and Brian Van Essen. 2016 · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang. 2016 · 2016
Earlier work this paper cites.
QSGD: Communication-efficient SGD via gradient quantization and encoding
Dan Alistarh, Demjan Grubic, Jerry Li, Ryota Tomioka, and Milan Vojnovic. 2017 · 2017
Earlier work this paper cites.
NVIDIA DGX-1
dgx1 2017 · 2017
Earlier work this paper cites.
NVIDIA Collective Communications Library (NCCL)
N NVIDIA. 2017 · 2017
Earlier work this paper cites.
NVIDIA NVLINK
NVLink 2017 · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
NVIDIA DGX-2
dgx2 2018 · 2018
Cited alongside, same era.
Pipedream: Fast and efficient pipeline parallel dnn training
Aaron Harlap, Deepak Narayanan, Amar Phanishayee, Vivek Seshadri, Nikhil Devanur, Greg Ganger, and Phil Gibbons. 2018 · 2018
Cited alongside, same era.
GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism
Yanping Huang, Yonglong Cheng, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, and Zhifeng Chen. 2018 · 2018
Cited alongside, same era.
NVIDIA NVSWITCH
NVSwitch 2018 · 2018
Cited alongside, same era.
Zero: Memory optimizations toward training trillion parameter models. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 1–16
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020 · 2020
Later among the works it cites.
DeepSpeed: Extreme-scale model training for everyone
DeepSpeed Team and Rangan Majumder. 2020 · 2020
Later among the works it cites.
Blink: Fast and Generic Collectives for Distributed ML. In Proceedings of Machine Learning and Systems 2020, MLSys 2020, Austin, TX, USA, March 2-4, 2020 , Inderjit S. Dhillon, Dimitris S. Papailiopoulos, and Vivienne Sze (Eds.). mlsys.org
Guanhua Wang, Shivaram Venkataraman, Amar Phanishayee, Jorgen Thelin, Nikhil R. Devanur, and Ion Stoica. 2020 · 2020
Later among the works it cites.
NVIDIA InfiniBand Adaptive Routing Technology
Infiniband Sharp white paper 2021 · 2021
Later among the works it cites.
1-bit LAMB: Communication Efficient Large-Scale Large-Batch Training with LAMB’s Convergence Speed
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Demystifying parallel and distributed deep learning: An in-depth concurrency analysis
Tal Ben-Nun and Torsten Hoefler. 2019 · 2019
Cited alongside, same era.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, et al · 2019
Cited alongside, same era.
PipeDream: Generalized Pipeline Parallelism for DNN Training. In ACM Symposium on Operating Systems Principles (SOSP 2019)
Deepak Narayanan, Aaron Harlap, Amar Phanishayee, Vivek Seshadri, Nikhil Devanur, Greg Granger, Phil Gibbons, and Matei Zaharia. 2019 · 2019
Cited alongside, same era.
NVIDIA TESLA V100 GPU ACCELERATOR
Nvidia V100 datasheet 2017 · 2019
Cited alongside, same era.
Language Models are Unsupervised Multitask Learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Cited alongside, same era.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019 · 2019
Cited alongside, same era.
Improving Neural Network Quantization without Retraining using Outlier Channel Splitting. In Proceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA (Proceedings of Machine Learning Research, Vol. 97) , Kamalika Chaudhuri and Ruslan Salakhutdinov (Eds.). PMLR, 7543–7552
Ritchie Zhao, Yuwei Hu, Jordan Dotzel, Christopher De Sa, and Zhiru Zhang. 2019 · 2019
Cited alongside, same era.
Conglong Li, Ammar Ahmad Awan, Hanlin Tang, Samyam Rajbhandari, and Yuxiong He. 2021 · 2021
Later among the works it cites.
ZeRO-Infinity: Breaking the GPU Memory Wall for Extreme Scale Deep Learning. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC ’21)
Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, and Yuxiong He. 2021 · 2021
Later among the works it cites.
1-bit Adam: Communication Efficient Large-Scale Training with Adam’s Convergence Speed
Hanlin Tang, Shaoduo Gan, Ammar Ahmad Awan, Samyam Rajbhandari, Conglong Li, Xiangru Lian, Ji Liu, Ce Zhang, and Yuxiong He. 2021 · 2021
Later among the works it cites.
Stella Biderman, Kieran Bicheno, and Leo Gao. 2022 · 2022
Later among the works it cites.
GPT-NeoX-20B: An Open-Source Autoregressive Language Model
Sid Black, Stella Biderman, Eric Hallahan, Quentin Anthony, Leo Gao, Laurence Golding, Horace He, Connor Leahy, Kyle McDonell, Jason Phang, Michael Pieler, USVSN Sai Prashanth, Shivanshu Purohit, Laria Reynolds, Jonathan Tow, Ben Wang, and Samuel Weinbach. 2022 · 2022
Later among the works it cites.
8-bit Optimizers via Block-wise Quantization. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022 . OpenReview.net
Tim Dettmers, Mike Lewis, Sam Shleifer, and Luke Zettlemoyer. 2022 · 2022
Later among the works it cites.
Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, et al · 2022
Later among the works it cites.
MiCS: Near-linear Scaling for Training Gigantic Model on Public Cloud
Zhen Zhang, Shuai Zheng, Yida Wang, Justin Chiu, George Karypis, Trishul Chilimbi, Mu Li, and Xin Jin. 2022 · 2022
Later among the works it cites.
Quantization - PyTorch documentation
Quantization - PyTorch documentation 2023 · 2023
Closest in time.