Fetching the paper…
Reading the bibliography…
Foundation models are becoming the dominant deep learning technologies.
A. Gokaslan and V. Cohen, “Openwebtext corpus,” http://Skylion007.github.io/OpenWebTextCorpus , 2019
2008
Earlier work this paper cites.
P. Patarasuk and X. Yuan, “Bandwidth optimal all-reduce algorithms for clusters of workstations,” Journal of Parallel and Distributed Computing , vol. 69, no. 2, pp. 117–124, 2009
2009
Earlier work this paper cites.
2016
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala, “Pytorch: An imperative style, high-performance deep learning library,” in Advances in Neural Information Processing Systems 32 , H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett, Eds. Curran Associates, Inc., 2019, pp. 8024–8035
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al. , “Language models are unsupervised multitask learners,” OpenAI blog , vol. 1, no. 8, p. 9, 2019
2019
Earlier work this paper cites.
D. Narayanan, A. Harlap, A. Phanishayee, V. Seshadri, N. R. Devanur, G. R. Ganger, P. B. Gibbons, and M. Zaharia, “Pipedream: generalized pipeline parallelism for dnn training,” in Proceedings of the 27th ACM Symposium on Operating Systems Principles , 2019, pp. 1–15
2019
Earlier work this paper cites.
Y. Huang, Y. Cheng, A. Bapna, O. Firat, D. Chen, M. Chen, H. Lee, J. Ngiam, Q. V. Le, Y. Wu et al. , “Gpipe: Efficient training of giant neural networks using pipeline parallelism,” Advances in neural information processing systems , vol. 32, 2019
2019
Earlier work this paper cites.
J. Zhan and J. Zhang, “Pipe-torch: Pipeline-based distributed deep learning in a gpu cluster with heterogeneous networking,” in 2019 Seventh International Conference on Advanced Cloud and Big Data (CBD) . IEEE, 2019, pp. 55–60
2019
Earlier work this paper cites.
“Nvidia collective communications library (nccl),” 2019. [Online]. Available: https://developer.nvidia.com/nccl
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
Microsoft, “Deepspeed: Extreme-scale model training for everyone,” https://www.microsoft.com/en-us/research/blog/deepspeed-extreme-scale-model-training-for-everyone/ , 2020
2020
Earlier work this paper cites.
T. Wolf, L. Debut, V. Sanh, J. Chaumond, C. Delangue, A. Moi, P. Cistac, C. Ma, Y. Jernite, J. Plu, C. Xu, T. Le Scao, S. Gugger, M. Drame, Q. Lhoest, and A. M. Rush, “Transformers: State-of-the-Art Natural Language Processing.” Association for Computational Linguistics, 10 2020, pp. 38–45. [Online]. Available: https://www.aclweb.org/anthology/2020.emnlp-demos.6
2020
Earlier work this paper cites.
J. Rasley, S. Rajbhandari, O. Ruwase, and Y. He, DeepSpeed: System Optimizations Enable Training Deep Learning Models with Over 100 Billion Parameters . New York, NY, USA: Association for Computing Machinery, 2020, p. 3505–3506. [Online]. Available: https://doi.org/10.1145/3394486.3406703
2020
Earlier work this paper cites.
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu, “Exploring the limits of transfer learning with a unified text-to-text transformer,” 2020
2020
Earlier work this paper cites.
S. Li, Y. Zhao, R. Varma, O. Salpekar, P. Noordhuis, T. Li, A. Paszke, J. Smith, B. Vaughan, P. Damania, and S. Chintala, “Pytorch distributed: Experiences on accelerating data parallel training,” Proc. VLDB Endow. , vol. 13, no. 12, p. 3005–3018, aug 2020. [Online]. Available: https://doi.org/10.14778/3415478.3415530
2020
Cited alongside, same era.
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He, “Zero: Memory optimizations toward training trillion parameter models,” in SC20: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 2020, pp. 1–16
2020
Cited alongside, same era.
J. H. Park, G. Yun, C. M. Yi, N. T. Nguyen, S. Lee, J. Choi, S. H. Noh, and Y. ri Choi, “HetPipe: Enabling large DNN training on (whimpy) heterogeneous GPU clusters through integration of pipelined model parallelism and data parallelism,” in 2020 USENIX Annual Technical Conference (USENIX ATC 20) . USENIX Association, Jul. 2020, pp. 307–321. [Online]. Available: https://www.usenix.org/conference/atc20/presentation/park
2020
Cited alongside, same era.
S. Eliad, I. Hakimi, A. De Jagger, M. Silberstein, and A. Schuster, “Fine-tuning giant neural networks on commodity hardware with automatic pipeline model parallelism,” in 2021 USENIX Annual Technical Conference (USENIX ATC 21) , 2021, pp. 381–396
2021
Later among the works it cites.
A. Kosson, V. Chiley, A. Venigalla, J. Hestness, and U. Koster, “Pipelined backpropagation at scale: training large models without batches,” Proceedings of Machine Learning and Systems , vol. 3, pp. 479–501, 2021
2021
Later among the works it cites.
B. Yang, J. Zhang, J. Li, C. Ré, C. Aberger, and C. De Sa, “Pipemare: Asynchronous pipeline parallel dnn training,” Proceedings of Machine Learning and Systems , vol. 3, pp. 269–296, 2021
2021
Later among the works it cites.
S. Fan, Y. Rong, C. Meng, Z. Cao, S. Wang, Z. Zheng, C. Wu, G. Long, J. Yang, L. Xia et al. , “Dapple: A pipelined data parallel approach for training large models,” in Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming , 2021, pp. 431–445
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
P. Jain, A. Jain, A. Nrusimha, A. Gholami, P. Abbeel, J. Gonzalez, K. Keutzer, and I. Stoica, “Checkmate: Breaking the memory wall with optimal tensor rematerialization,” Proceedings of Machine Learning and Systems , vol. 2, pp. 497–511, 2020
2020
Cited alongside, same era.
2020
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
D. Narayanan, M. Shoeybi, J. Casper, P. LeGresley, M. Patwary, V. Korthikanti, D. Vainbrand, P. Kashinkunti, J. Bernauer, B. Catanzaro et al. , “Efficient large-scale language model training on gpu clusters using megatron-lm,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , 2021, pp. 1–15
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Cited alongside, same era.
2021
Later among the works it cites.
X. Ye, Z. Lai, S. Li, L. Cai, D. Sun, L. Qiao, and D. Li, “Hippie: A data-paralleled pipeline approach to improve memory-efficiency and scalability for large dnn training,” in 50th International Conference on Parallel Processing , 2021, pp. 1–10
2021
Later among the works it cites.
S. Li and T. Hoefler, “Chimera: efficiently training large-scale neural networks with bidirectional pipelines,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , 2021, pp. 1–14
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
2021
Later among the works it cites.
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , 2021
2021
Later among the works it cites.
2022
Closest in time.
S. Athlur, N. Saran, M. Sivathanu, R. Ramjee, and N. Kwatra, “Varuna: scalable, low-cost training of massive deep learning models,” in Proceedings of the Seventeenth European Conference on Computer Systems , 2022, pp. 472–487
2022
Closest in time.
J. Yuan, X. Li, C. Cheng, J. Liu, R. Guo, S. Cai, C. Yao, F. Yang, X. Yi, C. Wu, H. Zhang, and J. Zhao, “Oneflow: Redesign the distributed deep learning framework from scratch,” 2022
2022
Closest in time.
X. Jia, L. Jiang, A. Wang, W. Xiao, Z. Shi, J. Zhang, X. Li, L. Chen, Y. Li, Z. Zheng, X. Liu, and W. Lin, “Whale: Efficient giant model training over heterogeneous GPUs,” in 2022 USENIX Annual Technical Conference (USENIX ATC 22) . Carlsbad, CA: USENIX Association, Jul. 2022, pp. 673–688. [Online]. Available: https://www.usenix.org/conference/atc22/presentation/jia-xianyan
2022
Closest in time.
P. Liang, Y. Tang, X. Zhang, Y. Bai, T. Su, linbo qiao, Z. Lai, and D. Li, “A Survey on Auto-Parallelism of Neural Networks Training,” 4 2022. [Online]. Available: https://www.techrxiv.org/articles/preprint/A_Survey_on_Auto-Parallelism_of_Neural_Networks_Training/19522414
2022
Closest in time.
J. Reed, Z. DeVito, H. He, A. Ussery, and J. Ansel, “torch.fx: Practical program capture and transformation for deep learning in python,” Proceedings of Machine Learning and Systems , vol. 4, pp. 638–651, 2022
2022
Closest in time.
W. Liu, Z. Lai, S. Li, Y. Duan, K. Ge, and D. Li, “Autopipe: A fast pipeline parallelism approach with balanced partitioning and micro-batch slicing,” in 2022 IEEE International Conference on Cluster Computing (CLUSTER) . IEEE, 2022, pp. 301–312
2022
Closest in time.