Fetching the paper…
Reading the bibliography…
LLMs have seen rapid adoption in all domains.
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2020 · 1909
Earlier work this paper cites.
ZeRO: Memory Optimizations Toward Training Trillion Parameter Models
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020 · 1910
Earlier work this paper cites.
Lustre: Building a file system for 1000-node clusters. In Proceedings of the 2003 Linux symposium , Vol. 2003. Linux symposium, Ontario, Canada, 380–386
Philip Schwan et al · 2003
Earlier work this paper cites.
Berkeley lab checkpoint/restart (blcr) for linux clusters
Paul H Hargrove and Jason C Duell. 2006 · 2006
Earlier work this paper cites.
DMTCP: Transparent checkpointing for cluster computations and the desktop. In IPDPS’09: International Symposium on Parallel & Distributed Processing . IEEE, Rome, Italy, 1–12
Jason Ansel, Kapil Arya, and Gene Cooperman. 2009 · 2009
Earlier work this paper cites.
CheCUDA: A Checkpoint/Restart Tool for CUDA Applications. In PDCAT’09: The International Conference on Parallel and Distributed Computing, Applications and Technologies . IEEE, Higashi-Hiroshima, Japan, 408–413
Hiroyuki Takizawa, Katsuto Sato, Kazuhiko Komatsu, and Hiroaki Kobayashi. 2009 · 2009
Earlier work this paper cites.
FTI: High performance Fault Tolerance Interface for hybrid systems. In SC’11: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, Seattle, WA, USA, 1–12
Leonardo Bautista-Gomez, Seiji Tsuboi, Dimitri Komatitsch, Franck Cappello, Naoya Maruyama, and Satoshi Matsuoka. 2011 · 2011
Earlier work this paper cites.
NVCR: A transparent checkpoint-restart library for NVIDIA CUDA. In IPDPS’11: Proceedings of the International Symposium on Parallel and Distributed Processing Workshops and Phd Forum . IEEE, Anchorage, AK, USA, 104–113
Akira Nukada, Hiroyuki Takizawa, and Satoshi Matsuoka. 2011 · 2011
Earlier work this paper cites.
Efficient inter-node MPI communication using GPUDirect RDMA for InfiniBand clusters with NVIDIA GPUs. In ICPP’13: The International Conference on Parallel Processing . IEEE, Lyon, France, 80–89
Sreeram Potluri, Khaled Hamidouche, Akshay Venkatesh, Devendar Bureddy, and Dhabaleswar K Panda. 2013 · 2013
Earlier work this paper cites.
Very Deep Convolutional Networks for Large-Scale Image Recognition
Karen Simonyan and Andrew Zisserman. 2015 · 2015
Earlier work this paper cites.
Deep Residual Learning for Image Recognition. In Conference on Computer Vision and Pattern Recognition (CVPR) . IEEE, Las Vegas, USA, 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Diederik P. Kingma and Jimmy Ba. 2017 · 2017
Earlier work this paper cites.
An overview of gradient descent optimization algorithms
Sebastian Ruder. 2017 · 2017
Earlier work this paper cites.
GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism. In NeurIPS’19: Advances in Neural Information Processing Systems , H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett (Eds.), Vol. 32. Curran Associates, Inc., Vancouver, Canada
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, and zhifeng Chen. 2019 · 2019
Earlier work this paper cites.
VeloC: Towards High Performance Adaptive Asynchronous Checkpointing at Large Scale. In IPDPS’19: IEEE International Parallel and Distributed Processing Symposium . IEEE, Rio de Janeiro, Brazil, 911–920
Bogdan Nicolae, Adam Moody, Elsa Gonsiorowski, Kathryn Mohror, and Franck Cappello. 2019 · 2019
Cited alongside, same era.
Adios 2: The adaptable input output system. a framework for high-performance data management
William F Godoy, Norbert Podhorszki, Ruonan Wang, Chuck Atkins, Greg Eisenhauer, Junmin Gu, Philip Davis, Jong Choi, Kai Germaschewski, Kevin Huck, et al · 2020
Cited alongside, same era.
Paxos vs Raft: have we reached consensus on distributed consensus?. In PaPoC’20: The 7th Workshop on Principles and Practice of Consistency for Distributed Data . ACM, Heraklion, Greece, Article 8, 9 pages
Heidi Howard and Richard Mortier. 2020 · 2020
Cited alongside, same era.
PyTorch Distributed: Experiences on Accelerating Data Parallel Training
Shen Li, Yanli Zhao, Rohan Varma, Omkar Salpekar, Pieter Noordhuis, Teng Li, Adam Paszke, Jeff Smith, Brian Vaughan, Pritam Damania, and Soumith Chintala. 2020 · 2020
Cited alongside, same era.
Towards Efficient Cache Allocation for High-Frequency Checkpointing. In HiPC’22: The 29th IEEE International Conference on High Performance Computing, Data, and Analytics . IEEE, Bangalore, India, 262–271
Avinash Maurya, Bogdan Nicolae, M. Mustafa Rafique, Amr M. Elsayed, Thierry Tonellot, and Franck Cappello. 2022 · 2022
Later among the works it cites.
A Cost-Efficient Failure-Tolerant Scheme for Distributed DNN Training. In ICCD’23: Proceedings of the International Conference on Computer Design . IEEE, Milan, Italy, 150–157
Menglei Chen, Yu Hua, Rong Bai, and Jianming Huang. 2023 · 2023
Later among the works it cites.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2023
Later among the works it cites.
Modeling Multi-Threaded Aggregated I/O for Asynchronous Checkpointing on HPC Systems. In ISPDC’23: Proceedings of the International Conference on Parallel and Distributed Computing . IEEE, Bucharest, Romania, 101–105
Mikaila Gossman, Bogdan Nicolae, and Jon Calhoun. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
DeepFreeze: Towards Scalable Asynchronous Checkpointing of Deep Learning Models. In CCGrid’20: The 20th International Symposium on Cluster, Cloud and Internet Computing . IEEE/ACM, Melbourne, Australia, 172–181
Bogdan Nicolae, Jiali Li, Justin M. Wozniak, George Bosilca, Matthieu Dorier, and Franck Cappello. 2020 · 2020
Cited alongside, same era.
Checkpoint restart support for heterogeneous hpc applications. In CCGRID’20: The International Symposium on Cluster, Cloud and Internet Computing (CCGRID) . IEEE/ACM, Melbourne, Australia, 242–251
Konstantinos Parasyris, Kai Keller, Leonardo Bautista-Gomez, and Osman Unsal. 2020 · 2020
Cited alongside, same era.
DeepSpeed: System Optimizations Enable Training Deep Learning Models with Over 100 Billion Parameters. In KDD’20: The 26th SIGKDD International Conference on Knowledge Discovery & Data Mining . ACM, Virtual Event CA USA, 3505–3506
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020 · 2020
Cited alongside, same era.
NVIDIA A100 Tensor Core GPU: Performance and Innovation
Jack Choquette, Wishwesh Gandhi, Olivier Giroux, Nick Stam, and Ronny Krashinsky. 2021 · 2021
Cited alongside, same era.
Towards Efficient I/O Scheduling for Collaborative Multi-Level Checkpointing. In MASCOTS’21: The 29th IEEE International Symposium on the Modeling, Analysis, and Simulation of Computer and Telecommunication Systems . IEEE, Virtual, Portugal, 1–8
Avinash Maurya, Bogdan Nicolae, Mustafa Rafique, Thierry Tonellot, and Franck Cappello. 2021 · 2021
Cited alongside, same era.
CheckFreq: Frequent, Fine-Grained DNN Checkpointing. In FAST’21: The 19th USENIX Conference on File and Storage Technologies . USENIX Association, Boston, USA, 203–216
Jayashree Mohan, Amar Phanishayee, and Vijay Chidambaram. 2021 · 2021
Cited alongside, same era.
ZeRO-infinity: breaking the GPU memory wall for extreme scale deep learning. In SC’21: The International Conference for High Performance Computing, Networking, Storage and Analysis . ACM, St. Louis, Missouri, Article 59, 14 pages
Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, and Yuxiong He. 2021 · 2021
Cited alongside, same era.
Canary: Fault-Tolerant FaaS for Stateful Time-Sensitive Applications. In SC22: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, Dallas, TX, USA, 1–16
Moiz Arif, Kevin Assogba, and M. Mustafa Rafique. 2022 · 2022
Cited alongside, same era.
Unicron: Economizing Self-Healing LLM Training at Scale
Tao He, Xue Li, Zhibin Wang, Kun Qian, Jingbo Xu, Wenyuan Yu, and Jingren Zhou. 2023 · 2023
Later among the works it cites.
Welcome to PyTorch Lightning — PyTorch Lightning 2.1.0 Documentation
PyTorch Lightning. 2023 · 2023
Later among the works it cites.
Optimize Checkpoint Performance for Large Models - Azure Machine Learning
Microsoft. 2023 · 2023
Later among the works it cites.
Shuaiwen Leon Song, Bonnie Kruft, Minjia Zhang, Conglong Li, Shiyang Chen, et al · 2023
Later among the works it cites.
Llama 2: Open Foundation and Fine-Tuned Chat Models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, et al · 2023
Later among the works it cites.
BLOOM: A 176B-Parameter Open-Access Multilingual Language Model
BigScience Workshop, Teven Le Scao, Angela Fan, Christopher Akiki, Ellie Pavlick, Suzana Ilić, et al · 2023
Later among the works it cites.
TRANSOM: An Efficient Fault-Tolerant System for Training LLMs
Baodong Wu, Lei Xia, Qingping Li, Kangyu Li, Xu Chen, Yongqiang Guo, Tieyao Xiang, Yuheng Chen, and Shigang Li. 2023 · 2023
Later among the works it cites.
GLM-130B: An Open Bilingual Pre-trained Model
Aohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang, Hanyu Lai, Ming Ding, Zhuoyi Yang, Yifan Xu, Wendi Zheng, Xiao Xia, Weng Lam Tam, Zixuan Ma, Yufei Xue, Jidong Zhai, Wenguang Chen, Peng Zhang, Yuxiao Dong, and Jie Tang. 2023 · 2023
Later among the works it cites.
Welcome to the TorchSnapshot documentation
PyTorch. 2024 · 2024
Closest in time.
AsyncCheckpointIO– PyTorch Lightning
PyTorch-Lightning. 2024 · 2024
Closest in time.