Fetching the paper…
Reading the bibliography…
This paper presents a low-cost network architecture for training large language models (LLMs) at hyperscale.
S. Coll, E. Frachtenberg, F. Petrini, A. Hoisie, and L. Gurvits, “Using multirail networks in high-performance clusters,” in Proceedings 2001 IEEE International Conference on Cluster Computing , 2001, pp. 15–24
2001
Earlier work this paper cites.
M. Al-Fares, A. Loukissas, and A. Vahdat, “A scalable, commodity data center network architecture,” in Proceedings of the ACM SIGCOMM 2008 Conference on Data Communication , ser. SIGCOMM ’08. New York, NY, USA: Association for Computing Machinery, 2008, p. 63–74. [Online]. Available: https://doi.org/10.1145/1402958.1402967
2008
Earlier work this paper cites.
Meta, “Introducing data center fabric, the next-generation facebook data center network,” 2014. [Online]. Available: https://engineering.fb.com/2014/11/14/production-engineering/introducing-data-center-fabric-the-next-generation-facebook-data-center-network/
2014
Earlier work this paper cites.
C. Guo, H. Wu, Z. Deng, G. Soni, J. Ye, J. Padhye, and M. Lipshteyn, “Rdma over commodity ethernet at scale,” in Proceedings of the 2016 ACM SIGCOMM Conference , ser. SIGCOMM ’16. New York, NY, USA: Association for Computing Machinery, 2016, p. 202–215. [Online]. Available: https://doi.org/10.1145/2934872.2934908
2016
Earlier work this paper cites.
T. Schneider, O. Bibartiu, and T. Hoefler, “Ensuring deadlock-freedom in low-diameter infiniband networks,” in 2016 IEEE 24th Annual Symposium on High-Performance Interconnects (HOTI) , 2016, pp. 1–8
2016
Earlier work this paper cites.
S. Hu, Y. Zhu, P. Cheng, C. Guo, K. Tan, J. Padhye, and K. Chen, “Deadlocks in datacenter networks: Why do they form, and how to avoid them,” in Proceedings of the 15th ACM Workshop on Hot Topics in Networks , ser. HotNets ’16. New York, NY, USA: Association for Computing Machinery, 2016, p. 92–98. [Online]. Available: https://doi.org/10.1145/3005745.3005760
2016
Earlier work this paper cites.
W. M. Mellette, R. McGuinness, A. Roy, A. Forencich, G. Papen, A. C. Snoeren, and G. Porter, “Rotornet: A scalable, low-complexity, optical datacenter network,” in Proceedings of the Conference of the ACM Special Interest Group on Data Communication , ser. SIGCOMM ’17. New York, NY, USA: Association for Computing Machinery, 2017, p. 267–280. [Online]. Available: https://doi.org/10.1145/3098822.3098838
2017
Earlier work this paper cites.
W. Xiao, R. Bhardwaj, R. Ramjee, M. Sivathanu, N. Kwatra, Z. Han, P. Patel, X. Peng, H. Zhao, Q. Zhang, F. Yang, and L. Zhou, “Gandiva: Introspective cluster scheduling for deep learning,” in 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18) . Carlsbad, CA: USENIX Association, Oct. 2018, pp. 595–610. [Online]. Available: https://www.usenix.org/conference/osdi18/presentation/xiao
2018
Earlier work this paper cites.
J. Gu, M. Chowdhury, K. G. Shin, Y. Zhu, M. Jeon, J. Qian, H. Liu, and C. Guo, “Tiresias: A GPU cluster manager for distributed deep learning,” in 16th USENIX Symposium on Networked Systems Design and Implementation (NSDI 19) . Boston, MA: USENIX Association, Feb. 2019, pp. 485–500. [Online]. Available: https://www.usenix.org/conference/nsdi19/presentation/gu
2019
Earlier work this paper cites.
T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, “Language models are few-shot learners,” 2020
2020
Earlier work this paper cites.
L. Labs, “Openai’s gpt-3 language model: A technical overview,” 2020. [Online]. Available: https://lambdalabs.com/blog/demystifying-gpt-3
2020
Earlier work this paper cites.
W. M. Mellette, R. Das, Y. Guo, R. McGuinness, A. C. Snoeren, and G. Porter, “Expanding across time to deliver bandwidth efficiency and low latency,” in 17th USENIX Symposium on Networked Systems Design and Implementation (NSDI 20) . Santa Clara, CA: USENIX Association, Feb. 2020, pp. 1–18. [Online]. Available: https://www.usenix.org/conference/nsdi20/presentation/mellette
2020
Earlier work this paper cites.
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei, “Scaling laws for neural language models,” 2020
2020
Earlier work this paper cites.
H. Ballani, P. Costa, R. Behrendt, D. Cletheroe, I. Haller, K. Jozwik, F. Karinou, S. Lange, K. Shi, B. Thomsen, and H. Williams, “Sirius: A flat datacenter network with nanosecond optical switching,” in Proceedings of the Annual Conference of the ACM Special Interest Group on Data Communication on the Applications, Technologies, Architectures, and Protocols for Computer Communication , ser. SIGCOMM ’20. New York, NY, USA: Association for Computing Machinery, 2020, p. 782–797. [Online]. Available: https://doi.org/10.1145/3387514.3406221
2020
Earlier work this paper cites.
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro, “Megatron-lm: Training multi-billion parameter language models using model parallelism,” 2020
2020
Earlier work this paper cites.
D. Narayanan, M. Shoeybi, J. Casper, P. LeGresley, M. Patwary, V. A. Korthikanti, D. Vainbrand, P. Kashinkunti, J. Bernauer, B. Catanzaro, A. Phanishayee, and M. Zaharia, “Efficient large-scale language model training on gpu clusters using megatron-lm,” 2021
2021
Earlier work this paper cites.
P. Goyal, P. Shah, K. Zhao, G. Nikolaidis, M. Alizadeh, and T. E. Anderson, “Backpressure flow control,” in 19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22) . Renton, WA: USENIX Association, Apr. 2022, pp. 779–805. [Online]. Available: https://www.usenix.org/conference/nsdi22/presentation/goyal
2022
Cited alongside, same era.
——, “Doubling all2all performance with nvidia collective communication library 2.12,” 2022. [Online]. Available: https://developer.nvidia.com/blog/doubling-all2all-performance-with-nvidia-collective-communication-library-2-12/
2022
Cited alongside, same era.
S. Rajbhandari, C. Li, Z. Yao, M. Zhang, R. Y. Aminabadi, A. A. Awan, J. Rasley, and Y. He, “Deepspeed-moe: Advancing mixture-of-experts inference and training to power next-generation ai scale,” 2022
2022
Cited alongside, same era.
NVIDIA, “Dgx h100 computer,” 2023. [Online]. Available: https://www.nvidia.com/en-us/data-center/dgx-h100/
2023
Closest in time.
Nvidia, “Nvidia dgx gh200,” 2023. [Online]. Available: https://www.nvidia.com/en-us/data-center/dgx-gh200/
2023
Closest in time.
Nvidia, “Nvidia dgx superpod: Next generation scalable infrastructure for ai leadership, reference architecture,” 2023. [Online]. Available: https://docs.nvidia.com/dgx-superpod-reference-architecture-with-dgx-h100-systems.pdf
2023
Closest in time.
N. P. Jouppi, G. Kurian, S. Li, P. Ma, R. Nagarajan, L. Nai, N. Patil, S. Subramanian, A. Swing, B. Towles, C. Young, X. Zhou, Z. Zhou, and D. Patterson, “Tpu v4: An optically reconfigurable supercomputer for machine learning with hardware support for embeddings,” 2023
2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Poutievski, O. Mashayekhi, J. Ong, A. Singh, M. Tariq, R. Wang, J. Zhang, V. Beauregard, P. Conner, S. Gribble, R. Kapoor, S. Kratzer, N. Li, H. Liu, K. Nagaraj, J. Ornstein, S. Sawhney, R. Urata, L. Vicisano, K. Yasumura, S. Zhang, J. Zhou, and A. Vahdat, “Jupiter evolving: Transforming google’s datacenter network via optical circuit switches and software-defined networking,” in Proceedings of the ACM SIGCOMM 2022 Conference , ser. SIGCOMM ’22. New York, NY, USA: Association for Computing Machinery, 2022, p. 66–85. [Online]. Available: https://doi.org/10.1145/3544216.3544265
2022
Cited alongside, same era.
V. Korthikanti, J. Casper, S. Lym, L. McAfee, M. Andersch, M. Shoeybi, and B. Catanzaro, “Reducing activation recomputation in large transformer models,” 2022
2022
Cited alongside, same era.
R. Pope, S. Douglas, A. Chowdhery, J. Devlin, J. Bradbury, A. Levskaya, J. Heek, K. Xiao, S. Agrawal, and J. Dean, “Efficiently scaling transformer inference,” 2022
2022
Cited alongside, same era.
Y. Zhao, Y. Liu, Y. Peng, Y. Zhu, X. Liu, and X. Jin, “Multi-resource interleaving for deep learning training,” in Proceedings of the ACM SIGCOMM 2022 Conference , ser. SIGCOMM ’22. New York, NY, USA: Association for Computing Machinery, 2022, p. 428–440. [Online]. Available: https://doi.org/10.1145/3544216.3544224
2022
Cited alongside, same era.
L. Zheng, Z. Li, H. Zhang, Y. Zhuang, Z. Chen, Y. Huang, Y. Wang, Y. Xu, D. Zhuo, E. P. Xing, J. E. Gonzalez, and I. Stoica, “Alpa: Automating inter- and Intra-Operator parallelism for distributed deep learning,” in 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22) . Carlsbad, CA: USENIX Association, Jul. 2022, pp. 559–578. [Online]. Available: https://www.usenix.org/conference/osdi22/presentation/zheng-lianmin
2022
Cited alongside, same era.
C. Unger, Z. Jia, W. Wu, S. Lin, M. Baines, C. E. Q. Narvaez, V. Ramakrishnaiah, N. Prajapati, P. McCormick, J. Mohd-Yusof, X. Luo, D. Mudigere, J. Park, M. Smelyanskiy, and A. Aiken, “Unity: Accelerating DNN training through joint optimization of algebraic transformations and parallelization,” in 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22) . Carlsbad, CA: USENIX Association, Jul. 2022, pp. 267–284. [Online]. Available: https://www.usenix.org/conference/osdi22/presentation/unger
2022
Cited alongside, same era.
OpenAI, “Gpt-4 technical report,” 2023
2023
Cited alongside, same era.
The-Decoder, “Gpt-4 has a trillion parameters - report,” 2023. [Online]. Available: https://the-decoder.com/gpt-4-has-a-trillion-parameters/
2023
Cited alongside, same era.
Nvidia, “Nvlink and nvswitch: The building blocks of advanced multi-gpu communication—within and between servers.” 2023. [Online]. Available: https://www.nvidia.com/en-us/data-center/nvlink/
2023
Cited alongside, same era.
M. Isaev, N. Mcdonald, L. Dennison, and R. Vuduc, “Calculon: a methodology and tool for high-level co-design of systems and large language models,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis , ser. SC ’23. New York, NY, USA: Association for Computing Machinery, 2023. [Online]. Available: https://doi.org/10.1145/3581784.3607102
2023
Closest in time.
Nvidia, “Nvidia docs hub: Mma4z00-ns400 400gb/s single-port osfp 400gb/s multimode sr4 50m connectivity scenarios,” 2023. [Online]. Available: https://docs.nvidia.com/networking/display/mma4z00ns400/connectivity+scenarios#:~:text=The%20400Gb%2Fs%20transceiver%20has,maximum%20or%208%20Watts%20typical
2023
Closest in time.
——, “Nvidia docs hub: Qm9700/qm9790 1u ndr 400gb/s infiniband switch systems user manual specifications,” 2023. [Online]. Available: https://docs.nvidia.com/networking/display/qm97x0pub/specifications
2023
Closest in time.
OpenAI, “Openai: Ai and compute,” 2023. [Online]. Available: https://openai.com/research/ai-and-compute
2023
Closest in time.
Databricks, “Hello dolly: Democratizing the magic of chatgpt with open models,” 2023. [Online]. Available: https://www.databricks.com/blog/2023/03/24/hello-dolly-democratizing-magic-chatgpt-open-models.html
2023
Closest in time.
Microsoft, “Deepspeed zero++: A leap in speed for llm and chat model training with 4x less communication,” 2023. [Online]. Available: https://www.microsoft.com/en-us/research/blog/deepspeed-zero-a-leap-in-speed-for-llm-and-chat-model-training-with-4x-less-communication/
2023
Closest in time.
Z. Li, L. Zheng, Y. Zhong, V. Liu, Y. Sheng, X. Jin, Y. Huang, Z. Chen, H. Zhang, J. E. Gonzalez, and I. Stoica, “AlpaServe: Statistical multiplexing with model parallelism for deep learning serving,” in 17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23) . Boston, MA: USENIX Association, Jul. 2023. [Online]. Available: https://www.usenix.org/conference/osdi23/presentation/li-zhouhan
2023
Closest in time.
S. Rajasekaran, M. Ghobadi, and A. Akella, “Cassini: Network-aware job scheduling in machine learning clusters,” 2023
2023
Closest in time.
D. Mudigere, Y. Hao, J. Huang, Z. Jia, A. Tulloch, S. Sridharan, X. Liu, M. Ozdal, J. Nie, J. Park, L. Luo, J. A. Yang, L. Gao, D. Ivchenko, A. Basant, Y. Hu, J. Yang, E. K. Ardestani, X. Wang, R. Komuravelli, C.-H. Chu, S. Yilmaz, H. Li, J. Qian, Z. Feng, Y. Ma, J. Yang, E. Wen, H. Li, L. Yang, C. Sun, W. Zhao, D. Melts, K. Dhulipala, K. Kishore, T. Graf, A. Eisenman, K. K. Matam, A. Gangidi, G. J. Chen, M. Krishnan, A. Nayak, K. Nair, B. Muthiah, M. khorashadi, P. Bhattacharya, P. Lapukhov, M. Naumov, A. Mathews, L. Qiao, M. Smelyanskiy, B. Jia, and V. Rao, “Software-hardware co-design for fast and scalable training of deep learning recommendation models,” 2023
2023
Closest in time.
L. Zhao and A. Krishnamurthy, “Bandwidth optimal pipeline schedule for collective communication,” 2023
2023
Closest in time.
——, “Gb200 nvl72 computer,” 2024. [Online]. Available: https://www.nvidia.com/en-us/data-center/gb200-nvl72/
2024
Closest in time.