Fetching the paper…
Reading the bibliography…
AI Infrastructure plays a key role in the speed and cost-competitiveness of developing and deploying advanced AI models.
A first order approximation to the optimum checkpoint interval
John W. Young · 1974
Earlier work this paper cites.
Gpfs: A shared-disk file system for large computing clusters
Frank Schmuck and Roger Haskin · 2002
Earlier work this paper cites.
Sr-iov: Performance benefits for virtualized interconnects
Glenn K Lockwood, Mahidhar Tatineni, and Rick Wagner · 2014
Earlier work this paper cites.
Scaling neural machine translation, 2018
Myle Ott, Sergey Edunov, David Grangier, and Michael Auli · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding, 2019
Devlin Jacob, Chang Ming-Wei, Lee Kenton, and Toutanova Kristina · 2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, et al · 2019
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer, 2020
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu · 2020
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism, 2020
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2020
Cited alongside, same era.
The high-speed networks of the summit and sierra supercomputers
C. B. Stunkel, R. L. Graham, G. Shainer, M. Kagan, S. S. Sharkawi, B. Rosenburg, and G. A. Chochia · 2020
Cited alongside, same era.
Efficient large-scale language model training on gpu clusters using megatron-lm
Deepak Narayanan, Mohammad Shoeybi, Jared Casper, et al · 2021
Cited alongside, same era.
How to deploy a high-performance distributed ai training cluster with nvidia gpus and kvm, 2022
Apoorve Mohan and Matthew Sheard · 2022
Cited alongside, same era.
Best practices for hpc workloads on public cloud platforms: A guide for computational scientists to use public cloud for hpc workloads
Robert Walkup, Seetharami R Seelam, and Sophia Wen · 2022
Cited alongside, same era.
To virtualize or not to virtualize ai infrastructure: A perspective
Seetharami Seelam, Apoorve Mohan, Ming-Hung Chen, and IHsin Chung · 2023
Later among the works it cites.
Bloomberggpt: A large language model for finance, 2023
Shijie Wu, Ozan Irsoy, Steven Lu, Vadim Dabravolski, Mark Dredze, Sebastian Gehrmann, Prabhanjan Kambadur, David Rosenberg, and Gideon Mann · 2023
Later among the works it cites.
URL https://engineering.fb.com/2024/03/12/data-center-engineering/building-metas-genai-infrastructure/
Building meta’s genai infrastructure · 2024
Closest in time.
MegaScale: Scaling large language model training to more than 10,000 GPUs
Ziheng Jiang, Haibin Lin, Yinmin Zhong, Qi Huang, Yangrui Chen, Zhi Zhang, Yanghua Peng, Xiang Li, Cong Xie, Shibiao Nong, Yulu Jia, Sun He, Hongmin Chen, Zhihao Bai, Qi Hou, Shipeng Yan, Ding Zhou, Yiyao Sheng, Zhuo Jiang, Haohan Xu, Haoran Wei, Zhang Zhang, Pengfei Nie, Leqi Zou, Sida Zhao, Liang Xiang, Zherui Liu, Zhe Li, Xiaoying Jia, Jianxi Ye, Xin Jin, and Xin Liu · 2024
Closest in time.
Granite code models: A family of open foundation models for code intelligence
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nvidia cuda samples, 2023
NVIDIA · 2023
Cited alongside, same era.
URL https://github.com/project-codeflare/multi-cluster-app-dispatcher
Multi-cluster app dispatcher, a
Cited in the paper.
URL https://www.ibm.com/artificial-intelligence
Ibm watsonx, b
Cited in the paper.
URL https://ai.meta.com/blog/meta-llama-3/
Introducing meta llama 3: The most capable openly available llm to date
Cited in the paper.
Supercharging ibm’s cloud-native ai supercomputer, 2023a
Talia Gershon, Bengi Karacali-Akyamac, Seetharami Seelam, Drew Thorstensen, and Rohit Badlaney
Cited in the paper.
Why we built an ai supercomputer in the cloud, 2023b
Talia Gershon, Seetharami Seelam, Jay Jubran, Eran Gampel, and Drew Thorstensen
Cited in the paper.
Ai training autopilot, 2024a
IBM Research
Cited in the paper.
Mayank Mishra, Matt Stallone, Gaoyuan Zhang, Yikang Shen, Aditya Prasad, Adriana Meza Soria, Michele Merler, Parameswaran Selvam, Saptha Surendran, Shivdeep Singh, Manish Sethi, Xuan-Hong Dang, Pengyuan Li, Kun-Lung Wu, Syed Zawad, Andrew Coleman, Matthew White, Mark Lewis, Raju Pavuluri, Yan Koyfman, Boris Lublinsky, Maximilien de Bayser, Ibrahim Abdelaziz, Kinjal Basu, Mayank Agarwal, Yi Zhou, Chris Johnson, Aanchal Goyal, Hima Patel, Yousaf Shah, Petros Zerfos, Heiko Ludwig, Asim Munawar, Maxwell Crouse, Pavan Kapanipathi, Brian Belgodere, Shweta Salaria, Bob Calio, Sophia Wen, Seetharami Seelam, Carlos Fonseca, Amith Singhee, Nirmit Desai, David D. Cox, Ruchir Puri, and Rameswar Panda · 2024
Closest in time.
Nvidia gpu memory error management, 2024
NVIDIA · 2024
Closest in time.