Fetching the paper…
Reading the bibliography…
Deep learning is experiencing a rise in large-scale models.
R. A. Van De Geijn and J. Watts, “SUMMA: Scalable universal matrix multiplication algorithm,” Concurrency: Practice and Experience , vol. 9, no. 4, pp. 255–274, 1997
1997
Earlier work this paper cites.
R. Thakur, R. Rabenseifner, and W. Gropp, “Optimization of collective communication operations in mpich,” The International Journal of High Performance Computing Applications , vol. 19, no. 1, pp. 49–66, 2005
2005
Earlier work this paper cites.
J. Forrest and R. Lougee-Heimer, “Cbc user guide,” in Emerging theory, methods, and applications . INFORMS, 2005, pp. 257–277
2005
Earlier work this paper cites.
P. Patarasuk and X. Yuan, “Bandwidth optimal all-reduce algorithms for clusters of workstations,” Journal of Parallel and Distributed Computing , vol. 69, no. 2, pp. 117–124, 2009
2009
Earlier work this paper cites.
J. Sanders and E. Kandrot, CUDA by example: an introduction to general-purpose GPU programming . Addison-Wesley Professional, 2010
2010
Earlier work this paper cites.
E. Georganas, J. González-Domínguez, E. Solomonik, Y. Zheng, J. Tourino, and K. Yelick, “Communication avoiding and overlapping for numerical linear algebra,” in SC’12: Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis . IEEE, 2012, pp. 1–11
2012
Earlier work this paper cites.
2016
Earlier work this paper cites.
H. Zhang, Z. Zheng, S. Xu, W. Dai, Q. Ho, X. Liang, Z. Hu, J. Wei, P. Xie, and E. P. Xing, “Poseidon: An efficient communication architecture for distributed deep learning on gpu clusters.” in USENIX Annual Technical Conference , vol. 1, no. 1, 2017, pp. 1–2
2017
Earlier work this paper cites.
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Advances in neural information processing systems , vol. 30, 2017
2017
Earlier work this paper cites.
2018
Earlier work this paper cites.
2018
Earlier work this paper cites.
2019
Earlier work this paper cites.
S. Shi, X. Chu, and B. Li, “MG-WFBP: Efficient data communication for distributed synchronous SGD algorithms,” in IEEE INFOCOM 2019-IEEE Conference on Computer Communications . IEEE, 2019, pp. 172–180
2019
Earlier work this paper cites.
Y. Peng, Y. Zhu, Y. Chen, Y. Bao, B. Yi, C. Lan, C. Wu, and C. Guo, “A generic communication scheduler for distributed dnn training acceleration,” in Proceedings of the 27th ACM Symposium on Operating Systems Principles , 2019, pp. 16–29
2019
Earlier work this paper cites.
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga et al. , “Pytorch: An imperative style, high-performance deep learning library,” Advances in neural information processing systems , vol. 32, 2019
2019
Earlier work this paper cites.
A. Li, S. L. Song, J. Chen, J. Li, X. Liu, N. R. Tallent, and K. J. Barker, “Evaluating modern gpu interconnect: Pcie, nvlink, nv-sli, nvswitch and gpudirect,” IEEE Transactions on Parallel and Distributed Systems , vol. 31, no. 1, pp. 94–110, 2019
2019
Earlier work this paper cites.
Y. Huang, Y. Cheng, A. Bapna, O. Firat, D. Chen, M. Chen, H. Lee, J. Ngiam, Q. V. Le, Y. Wu et al. , “Gpipe: Efficient training of giant neural networks using pipeline parallelism,” Advances in neural information processing systems , vol. 32, 2019
2019
Earlier work this paper cites.
D. Narayanan, A. Harlap, A. Phanishayee, V. Seshadri, N. R. Devanur, G. R. Ganger, P. B. Gibbons, and M. Zaharia, “PipeDream: Generalized pipeline parallelism for dnn training,” in Proceedings of the 27th ACM Symposium on Operating Systems Principles , 2019, pp. 1–15
2019
Earlier work this paper cites.
Z. Jia, M. Zaharia, and A. Aiken, “Beyond data and model parallelism for deep neural networks.” Proceedings of Machine Learning and Systems , vol. 1, pp. 1–13, 2019
2019
Earlier work this paper cites.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell et al. , “Language models are few-shot learners,” Advances in neural information processing systems , vol. 33, pp. 1877–1901, 2020
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
2020
Earlier work this paper cites.
P. Jain, A. Jain, A. Nrusimha, A. Gholami, P. Abbeel, J. Gonzalez, K. Keutzer, and I. Stoica, “Checkmate: Breaking the memory wall with optimal tensor rematerialization,” Proceedings of Machine Learning and Systems , vol. 2, pp. 497–511, 2020
2020
Cited alongside, same era.
S. Li, Y. Zhao, R. Varma, O. Salpekar, P. Noordhuis, T. Li, A. Paszke, J. Smith, B. Vaughan, P. Damania et al. , “Pytorch distributed: experiences on accelerating data parallel training,” Proceedings of the VLDB Endowment , vol. 13, no. 12, pp. 3005–3018, 2020
2020
Cited alongside, same era.
S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He, “ZeRO: Memory optimizations toward training trillion parameter models,” in SC20: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 2020, pp. 1–16
2020
Cited alongside, same era.
Y. Zhuang, L. Zheng, Z. Li, E. Xing, Q. Ho, J. Gonzalez, I. Stoica, H. Zhang, and H. Zhao, “On optimizing the communication of model parallelism,” Proceedings of Machine Learning and Systems , vol. 5, 2023
2022
Later among the works it cites.
2022
Later among the works it cites.
X. Jia, L. Jiang, A. Wang, W. Xiao, Z. Shi, J. Zhang, X. Li, L. Chen, Y. Li, Z. Zheng, X. Liu, and W. Lin, “Whale: Efficient giant model training over heterogeneous GPUs,” in 2022 USENIX Annual Technical Conference (USENIX ATC 22) . Carlsbad, CA: USENIX Association, Jul. 2022, pp. 673–688. [Online]. Available: https://www.usenix.org/conference/atc22/presentation/jia-xianyan
2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2020
Cited alongside, same era.
2021
Cited alongside, same era.
S. Eliad, I. Hakimi, A. De Jagger, M. Silberstein, and A. Schuster, “Fine-tuning giant neural networks on commodity hardware with automatic pipeline model parallelism,” in 2021 USENIX Annual Technical Conference (USENIX ATC 21) , 2021, pp. 381–396
2021
Cited alongside, same era.
2021
Cited alongside, same era.
M. Kirisame, S. Lyubomirsky, A. Haan, J. Brennan, M. He, J. Roesch, T. Chen, and Z. Tatlock, “Dynamic tensor rematerialization,” in International Conference on Learning Representations , 2021. [Online]. Available: https://openreview.net/forum?id=Vfs_2RnOD0H
2021
Cited alongside, same era.
D. Narayanan, M. Shoeybi, J. Casper, P. LeGresley, M. Patwary, V. Korthikanti, D. Vainbrand, P. Kashinkunti, J. Bernauer, B. Catanzaro et al. , “Efficient large-scale language model training on gpu clusters using megatron-lm,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , 2021, pp. 1–15
2021
Cited alongside, same era.
2021
Cited alongside, same era.
S. Fan, Y. Rong, C. Meng, Z. Cao, S. Wang, Z. Zheng, C. Wu, G. Long, J. Yang, L. Xia et al. , “DAPPLE: A pipelined data parallel approach for training large models,” in Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming , 2021, pp. 431–445
2021
Cited alongside, same era.
Z. Li, S. Zhuang, S. Guo, D. Zhuo, H. Zhang, D. Song, and I. Stoica, “Terapipe: Token-level pipeline parallelism for training large-scale language models,” in International Conference on Machine Learning . PMLR, 2021, pp. 6543–6552
2021
Cited alongside, same era.
2022
Later among the works it cites.
S. Black, S. Biderman, E. Hallahan, Q. Anthony, L. Gao, L. Golding, H. He, C. Leahy, K. McDonell, J. Phang et al. , “Gpt-neox-20b: An open-source autoregressive language model,” in Proceedings of BigScience Episode# 5–Workshop on Challenges & Perspectives in Creating Large Language Models , 2022, pp. 95–136
2022
Later among the works it cites.
X. Sun, W. Wang, S. Qiu, R. Yang, S. Huang, J. Xu, and Z. Wang, “Stronghold: fast and affordable billion-scale deep learning model training,” in Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis , 2022, pp. 1–17
2022
Later among the works it cites.
W. Liu, Z. Lai, S. Li, Y. Duan, K. Ge, and D. Li, “Autopipe: A fast pipeline parallelism approach with balanced partitioning and micro-batch slicing,” in 2022 IEEE International Conference on Cluster Computing (CLUSTER) . IEEE, 2022, pp. 301–312
2022
Later among the works it cites.
B. Wang, Q. Xu, Z. Bian, and Y. You, “Tesseract: Parallelize the tensor parallelism efficiently,” in Proceedings of the 51st International Conference on Parallel Processing , 2022, pp. 1–11
2022
Later among the works it cites.
S. Li, Z. Lai, D. Li, Y. Zhang, X. Ye, and Y. Duan, “Embrace: Accelerating sparse communication for distributed training of deep neural networks,” in Proceedings of the 51st International Conference on Parallel Processing , 2022, pp. 1–11
2022
Later among the works it cites.
H. Oh, J. Lee, H. Kim, and J. Seo, “Out-of-order backprop: an effective scheduling technique for deep learning,” in Proceedings of the Seventeenth European Conference on Computer Systems , 2022, pp. 435–452
2022
Later among the works it cites.
Y. Feng, M. Xie, Z. Tian, S. Wang, Y. Lu, and J. Shu, “Mobius: Fine tuning large-scale models on commodity gpu servers,” in Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 , 2023, pp. 489–501
2023
Closest in time.
K. Mahajan, C.-H. Chu, S. Sridharan, and A. Akella, “Better together: Jointly optimizing ML collective scheduling and execution planning using SYNDICATE,” in 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23) . Boston, MA: USENIX Association, Apr. 2023, pp. 809–824. [Online]. Available: https://www.usenix.org/conference/nsdi23/presentation/mahajan
2023
Closest in time.
Z. Lai, S. Li, X. Tang, K. Ge, W. Liu, Y. Duan, L. Qiao, and D. Li, “Merak: An efficient distributed dnn training framework with automated 3d parallelism for giant foundation models,” IEEE Transactions on Parallel and Distributed Systems , vol. 34, no. 5, pp. 1466–1478, 2023
2023
Closest in time.
2023
Closest in time.
S. Li, H. Liu, Z. Bian, J. Fang, H. Huang, Y. Liu, B. Wang, and Y. You, “Colossal-ai: A unified deep learning system for large-scale parallel training,” in Proceedings of the 52nd International Conference on Parallel Processing , 2023, pp. 766–775
2023
Closest in time.
P. Liang, Y. Tang, X. Zhang, Y. Bai, T. Su, Z. Lai, L. Qiao, and D. Li, “A survey on auto-parallelism of large-scale deep learning training,” IEEE Transactions on Parallel and Distributed Systems , 2023
2023
Closest in time.
2023
Closest in time.
V. A. Korthikanti, J. Casper, S. Lym, L. McAfee, M. Andersch, M. Shoeybi, and B. Catanzaro, “Reducing activation recomputation in large transformer models,” Proceedings of Machine Learning and Systems , vol. 5, 2023
2023
Closest in time.
Y. Zhao, A. Gu, R. Varma, L. Luo, C.-C. Huang, M. Xu, L. Wright, H. Shojanazeri, M. Ott, S. Shleifer, A. Desmaison, C. Balioglu, P. Damania, B. Nguyen, G. Chauhan, Y. Hao, A. Mathews, and S. Li, “Pytorch fsdp: Experiences on scaling fully sharded data parallel,” Proceedings of the VLDB Endowment , vol. 16, no. 12, p. 3848–3860, sep 2023
2023
Closest in time.
Q. Zhou, H. Wang, X. Yu, C. Li, Y. Bai, F. Yan, and Y. Xu, “MPress: Democratizing billion-scale model training on multi-gpu servers via memory-saving inter-operator parallelism,” in 2023 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . IEEE, 2023, pp. 556–569
2023
Closest in time.
vast.ai, “Vast pricing,” May 2024. [Online]. Available: https://vast.ai/pricing
2024
Closest in time.