Fetching the paper…
Reading the bibliography…
Heavy communication, in particular, collective operations, can become a critical performance bottleneck in scaling the training of billion-parameter neural networks to large-scale parallel systems.
R. C. Agarwal, S. M. Balle, F. G. Gustavson, M. Joshi, and P. Palkar, “A three-dimensional approach to parallel matrix multiplication,” IBM Journal of Research and Development , vol. 39, no. 5, pp. 575–582, 1995
1995
Earlier work this paper cites.
2001
Earlier work this paper cites.
R. Thakur and W. D. Gropp, “Improving the performance of collective operations in mpich,” in Recent Advances in Parallel Virtual Machine and Message Passing Interface , J. Dongarra, D. Laforenza, and S. Orlando, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 2003, pp. 257–267
2003
Earlier work this paper cites.
R. Rabenseifner, “Optimization of collective reduction operations,” in Computational Science - ICCS 2004 , M. Bubak, G. D. van Albada, P. M. A. Sloot, and J. Dongarra, Eds. Berlin, Heidelberg: Springer Berlin Heidelberg, 2004, pp. 1–9
2004
Earlier work this paper cites.
2005
Earlier work this paper cites.
2015
Earlier work this paper cites.
2015
Earlier work this paper cites.
T. Chen, B. Xu, C. Zhang, and C. Guestrin, “Training deep nets with sublinear memory cost,” 2016
2016
Earlier work this paper cites.
2017
Earlier work this paper cites.
M. Belkin, D. Hsu, S. Ma, and S. Mandal, “Reconciling modern machine-learning practice and the classical bias-variance trade-off,” Proceedings of the National Academy of Sciences , vol. 116, no. 32, pp. 15 849–15 854, 2019. [Online]. Available: https://www.pnas.org/doi/abs/10.1073/pnas.1903070116
2019
Earlier work this paper cites.
N. Dryden, N. Maruyama, T. Moon, T. Benson, M. Snir, and B. Van Essen, “Channel and filter parallelism for large-scale cnn training,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , ser. SC ’19. New York, NY, USA: Association for Computing Machinery, 2019. [Online]. Available: https://doi.org/10.1145/3295500.3356207
2019
Earlier work this paper cites.
L. Jiao, F. Zhang, F. Liu, S. Yang, L. Li, Z. Feng, and R. Qu, “A survey of deep learning-based object detection,” IEEE Access , vol. 7, pp. 128 837–128 868, 2019. [Online]. Available: https://doi.org/10.1109%2Faccess.2019.2939201
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language models are unsupervised multitask learners,” 2019
2019
Earlier work this paper cites.
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro, “Megatron-lm: Training multi-billion parameter language models using model parallelism,” 2020
2020
Cited alongside, same era.
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He, “Zero: Memory optimizations toward training trillion parameter models,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , ser. SC ’20. IEEE Press, 2020
2020
Cited alongside, same era.
A. Tripathy, K. Yelick, and A. Buluç, “Reducing communication in graph neural network training,” in SC20: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 2020, pp. 1–14
2020
Cited alongside, same era.
S. Li, Y. Zhao, R. Varma, O. Salpekar, P. Noordhuis, T. Li, A. Paszke, J. Smith, B. Vaughan, P. Damania, and S. Chintala, “Pytorch distributed: Experiences on accelerating data parallel training,” Proc. VLDB Endow. , vol. 13, no. 12, p. 3005–3018, Aug. 2020. [Online]. Available: https://doi.org/10.14778/3415478.3415530
S. Wang, J. Wei, A. Sabne, A. Davis, B. Ilbeyi, B. Hechtman, D. Chen, K. S. Murthy, M. Maggioni, Q. Zhang, S. Kumar, T. Guo, Y. Xu, and Z. Zhou, “Overlap communication with dependent computation via decomposition in large deep learning models,” in Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1 , ser. ASPLOS 2023. New York, NY, USA: Association for Computing Machinery, 2022, p. 93–106. [Online]. Available: https://doi.org/10.1145/3567955.3567959
2022
Later among the works it cites.
2022
Later among the works it cites.
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
Microsoft, “3d parallelism with megatronlm and zero redundancy optimizer,” https://github.com/microsoft/DeepSpeedExamples/tree/master/Megatron-LM-v1.1.5-3D_parallelism
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
A. Nichol and P. Dhariwal, “Improved denoising diffusion probabilistic models,” 2021
2021
Cited alongside, same era.
2021
Cited alongside, same era.
S. Smith, M. Patwary, B. Norick, P. LeGresley, S. Rajbhandari, J. Casper, Z. Liu, S. Prabhumoye, G. Zerveas, V. Korthikanti, E. Zhang, R. Child, R. Y. Aminabadi, J. Bernauer, X. Song, M. Shoeybi, Y. He, M. Houston, S. Tiwary, and B. Catanzaro, “Using deepspeed and megatron to train megatron-turing nlg 530b, a large-scale generative language model,” 2022
2022
Cited alongside, same era.
B. Wang, Q. Xu, Z. Bian, and Y. You, “Tesseract: Parallelize the tensor parallelism efficiently,” in Proceedings of the 51st International Conference on Parallel Processing . ACM, aug 2022
2022
Cited alongside, same era.
S. Minaee, Y. Boykov, F. Porikli, A. Plaza, N. Kehtarnavaz, and D. Terzopoulos, “Image segmentation using deep learning: A survey,” IEEE Transactions on Pattern Analysis and Machine Intelligence , vol. 44, no. 7, pp. 3523–3542, 2022
2022
Later among the works it cites.
BigScience, “Bigscience large open-science open-access multilingual language model,” https://huggingface.co/bigscience/bloom
2022
Later among the works it cites.
S. Black, S. Biderman, E. Hallahan, Q. Anthony, L. Gao, L. Golding, H. He, C. Leahy, K. McDonell, J. Phang, M. Pieler, U. S. Prashanth, S. Purohit, L. Reynolds, J. Tow, B. Wang, and S. Weinbach, “GPT-NeoX-20B: An open-source autoregressive language model,” in Proceedings of BigScience Episode #5 – Workshop on Challenges & Perspectives in Creating Large Language Models , A. Fan, S. Ilic, T. Wolf, and M. Gallé, Eds. virtual+Dublin: Association for Computational Linguistics, May 2022, pp. 95–136. [Online]. Available: https://aclanthology.org/2022.bigscience-1.9
2022
Later among the works it cites.
Z. Lai, S. Li, X. Tang, K. Ge, W. Liu, Y. Duan, L. Qiao, and D. Li, “Merak: An efficient distributed dnn training framework with automated 3d parallelism for giant foundation models,” IEEE Transactions on Parallel and Distributed Systems , vol. 34, no. 5, pp. 1466–1478, 2023
2023
Closest in time.
2023
Closest in time.
2023
Closest in time.
A. Bhatele, R. Dhakal, A. Movsesyan, A. Ranjan, J. Marry, and O. Cankur, “Pipit: Enabling programmatic analysis of parallel execution traces,” 2023
2023
Closest in time.
L. AI, “Litgpt,” https://github.com/Lightning-AI/litgpt
2023
Closest in time.
J. Parmar, S. Prabhumoye, J. Jennings, M. Patwary, S. Subramanian, D. Su, C. Zhu, D. Narayanan, A. Jhunjhunwala, A. Dattagupta, V. Jawa, J. Liu, A. Mahabaleshwarkar, O. Nitski, A. Brundyn, J. Maki, M. Martinez, J. You, J. Kamalu, P. LeGresley, D. Fridman, J. Casper, A. Aithal, O. Kuchaiev, M. Shoeybi, J. Cohen, and B. Catanzaro, “Nemotron-4 15b technical report,” 2024
2024
Closest in time.