Fetching the paper…
Reading the bibliography…
We study a novel and important communication pattern in large-scale model-parallel deep learning (DL), which we call cross-mesh resharding.
Efficient all-to-all communication patterns in hypercube and mesh topologies
Scott, D. S · 1991
Earlier work this paper cites.
Interprocessor collective communication library (intercom)
Barnett, M., Shuler, L., van De Geijn, R., Gupta, S., Payne, D. G., and Watts, J · 1994
Earlier work this paper cites.
Collective communication: theory, practice, and experience
Chan, E., Heimlich, M., Purkayastha, A., and Van De Geijn, R · 2007
Earlier work this paper cites.
Bandwidth optimal all-reduce algorithms for clusters of workstations
Patarasuk, P. and Yuan, X · 2009
Earlier work this paper cites.
Introducing data center fabric, the next-generation facebook data center network, 2014
Andreyev, A · 2014
Earlier work this paper cites.
Beyond data and model parallelism for deep neural networks
Jia, Z., Zaharia, M., and Aiken, A · 2018
Earlier work this paper cites.
The nvidia collective communication library, 2018
NVIDIA · 2018
Earlier work this paper cites.
Horovod: fast and easy distributed deep learning in tensorflow
Sergeev, A. and Del Balso, M · 2018
Earlier work this paper cites.
Mesh-tensorflow: Deep learning for supercomputers
Shazeer, N., Cheng, Y., Parmar, N., Tran, D., Vaswani, A., Koanantakool, P., Hawkins, P., Lee, H., Hong, M., Young, C., et al · 2018
Earlier work this paper cites.
Blueconnect: Decomposing all-reduce for deep learning on heterogeneous network hierarchy
Cho, M., Finkler, U., Kung, D., and Hunter, H · 2019
Earlier work this paper cites.
Tictac: Accelerating distributed deep learning with communication scheduling
Hashemi, S. H., Abdu Jyothi, S., and Campbell, R · 2019
Earlier work this paper cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, D., Chen, M., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., et al · 2019
Earlier work this paper cites.
Priority-based parameter propagation for distributed dnn training
Jayarajan, A., Wei, J., Gibson, G., Fedorova, A., and Pekhimenko, G · 2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., and Catanzaro, B · 2019
Cited alongside, same era.
Supporting very large models using automatic dataflow graph partitioning
Wang, M., Huang, C.-c., and Li, J · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Cited alongside, same era.
A unified architecture for accelerating distributed dnn training in heterogeneous gpu/cpu clusters
Jiang, Y., Zhu, Y., Lan, C., Yi, B., Cui, Y., and Guo, C · 2020
Cited alongside, same era.
Gshard: Scaling giant models with conditional computation and automatic sharding
Lepikhin, D., Lee, H., Xu, Y., Chen, D., Firat, O., Huang, Y., Krikun, M., Shazeer, N., and Chen, Z · 2020
U-net transformer: Self and cross attention for medical image segmentation
Petit, O., Thome, N., Rambour, C., Themyr, L., Collins, T., and Soler, L · 2021
Later among the works it cites.
Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning
Rajbhandari, S., Ruwase, O., Rasley, J., Smith, S., and He, Y · 2021
Later among the works it cites.
Memory-efficient array redistribution through portable collective communication
Rink, N. A., Paszke, A., Vytiniotis, D., and Schmid, G. S · 2021
Later among the works it cites.
Synthesizing collective communication algorithms for heterogeneous networks with taccl
Shah, A., Chidambaram, V., Cowan, M., Maleki, S., Musuvathi, M., Mytkowicz, T., Nelson, J., Saarikivi, O., and Singh, R · 2021
Later among the works it cites.
Gspmd: general and scalable parallelization for ml computation graphs
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Pytorch distributed: Experiences on accelerating data parallel training
Li, S., Zhao, Y., Varma, R., Salpekar, O., Noordhuis, P., Li, T., Paszke, A., Smith, J., Vaughan, B., Damania, P., et al · 2020
Cited alongside, same era.
Plink: Discovering and exploiting locality for accelerated distributed training on the public cloud
Luo, L., West, P., Nelson, J., Krishnamurthy, A., and Ceze, L · 2020
Cited alongside, same era.
Zero: Memory optimizations toward training trillion parameter models
Rajbhandari, S., Rasley, J., Ruwase, O., and He, Y · 2020
Cited alongside, same era.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Rasley, J., Rajbhandari, S., Ruwase, O., and He, Y · 2020
Cited alongside, same era.
Blink: Fast and generic collectives for distributed ml
Wang, G., Venkataraman, S., Phanishayee, A., Devanur, N., Thelin, J., and Stoica, I · 2020
Cited alongside, same era.
Autosync: Learning to synchronize for data-parallel distributed deep learning
Zhang, H., Li, Y., Deng, Z., Liang, X., Carin, L., and Xing, E · 2020
Cited alongside, same era.
Synthesizing optimal collective algorithms
Cai, Z., Liu, Z., Maleki, S., Musuvathi, M., Mytkowicz, T., Nelson, J., and Saarikivi, O · 2021
Cited alongside, same era.
Xu, Y., Lee, H., Chen, D., Hechtman, B., Huang, Y., Joshi, R., Krikun, M., Lepikhin, D., Ly, A., Maggioni, M., et al · 2021
Later among the works it cites.
Hoplite: efficient and fault-tolerant collective communication for task-based distributed systems
Zhuang, S., Li, Z., Zhuo, D., Wang, S., Liang, E., Nishihara, R., Moritz, P., and Stoica, I · 2021
Later among the works it cites.
Pathways: Asynchronous distributed dataflow for ml
Barham, P., Chowdhery, A., Dean, J., Ghemawat, S., Hand, S., Hurt, D., Isard, M., Lim, H., Pang, R., Roy, S., et al · 2022
Closest in time.
Gc3: An optimizing compiler for gpu collective communication
Cowan, M., Maleki, S., Musuvathi, M., Saarikivi, O., and Xiong, Y · 2022
Closest in time.
Reducing activation recomputation in large transformer models
Korthikanti, V., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., and Catanzaro, B · 2022
Closest in time.
High-resolution image synthesis with latent diffusion models
Rombach, R., Blattmann, A., Lorenz, D., Esser, P., and Ommer, B · 2022
Closest in time.
Unity: Accelerating { \{ DNN } \} training through joint optimization of algebraic transformations and parallelization
Unger, C., Jia, Z., Wu, W., Lin, S., Baines, M., Narvaez, C. E. Q., Ramakrishnaiah, V., Prajapati, N., McCormick, P., Mohd-Yusof, J., et al · 2022
Closest in time.
Synthesizing optimal parallelism placement and reduction strategies on hierarchical systems for deep learning
Xie, N., Norman, T., Grewe, D., and Vytiniotis, D · 2022
Closest in time.
Alpa: Automating inter-and intra-operator parallelism for distributed deep learning
Zheng, L., Li, Z., Zhang, H., Zhuang, Y., Chen, Z., Huang, Y., Wang, Y., Xu, Y., Zhuo, D., Gonzalez, J. E., et al · 2022
Closest in time.