Fetching the paper…
Reading the bibliography…
Model checkpoints are critical Deep Learning (DL) artifacts that enable fault tolerance for training and downstream applications, such as inference.
Distributed snapshots: determining global states of distributed systems
K Mani Chandy and Leslie Lamport · 1985
Earlier work this paper cites.
Checkpointing and recovery roll back for distributed systems
R Koo and S Toueg · 1987
Earlier work this paper cites.
An overview of checkpointing in uniprocessor and distributedsystems, focusing on implementation and performance
James S. Plank · 1997
Earlier work this paper cites.
Imagenet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Dram errors in the wild: a large-scale field study
Bianca Schroeder, Eduardo Pinheiro, and Wolf-Dietrich Weber · 2009
Earlier work this paper cites.
Design, modeling, and evaluation of a scalable multi-level checkpointing system
Adam Moody, Greg Bronevetsky, Kathryn Mohror, and Bronis R. de Supinski · 2010
Earlier work this paper cites.
Optimization of multi-level checkpoint model for large scale hpc applications
Sheng Di, Mohamed Slim Bouguerra, Leonardo Bautista-gomez, and Franck Cappello · 2014
Earlier work this paper cites.
Lessons learned from the analysis of system failures at petascale: The case of blue waters
Catello Di Martino, Zbigniew Kalbarczyk, Ravishankar K Iyer, Fabio Baccanico, Joseph Fullop, and William Kramer · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
Understanding practical tradeoffs in hpc checkpoint-scheduling policies
Nosayba El-Sayed and Bianca Schroeder · 2016
Earlier work this paper cites.
Scalable i/o-aware job scheduling for burst buffer enabled hpc clusters
Stephen Herbein, Dong H Ahn, Don Lipari, Thomas RW Scogland, Marc Stearman, Mark Grondona, Jim Garlick, Becky Springmeyer, and Michela Taufer · 2016
Earlier work this paper cites.
Flash storage disaggregation
Ana Klimovic, Christos Kozyrakis, Eno Thereska, Binu John, and Sanjeev Kumar · 2016
Earlier work this paper cites.
Leveraging near data processing for high-performance checkpoint/restart
Abhinav Agrawal, Gabriel H Loh, and James Tuck · 2017
Earlier work this paper cites.
Spin: Seamless operating system integration of peer-to-peer dma between ssds and gpus
Shai Bergman, Tanya Brokhman, Tzachi Cohen, and Mark Silberstein · 2017
Earlier work this paper cites.
Failures in large scale systems: long-term measurement, analysis, and implications
Saurabh Gupta, Tirthak Patel, Christian Engelmann, and Devesh Tiwari · 2017
Earlier work this paper cites.
Spin: Seamless operating system integration of peer-to-peer dma between ssds and gpus
Shai Bergman, Tanya Brokhman, Tzachi Cohen, and Mark Silberstein · 2019
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Cognitive ssd: A deep learning engine for in-storage data retrieval
hengwen Liang, Ying Wang, Youyou Lu, Zhe Yang, Huawei Li, and Xiaowei Li · 2019
Cited alongside, same era.
Deep learning recommendation model for personalization and recommendation systems
Maxim Naumov, Dheevatsa Mudigere, Hao-Jun Michael Shi, Jianyu Huang, Narayanan Sundaraman, Jongsoo Park, Xiaodong Wang, Udit Gupta, Carole-Jean Wu, Alisson G. Azzolini, Dmytro Dzhulgakov, Andrey Mallevich, Ilia Cherniavskii, Yinghai Lu, Raghuraman Krishnamoorthi, Ansha Yu, Volodymyr Kondratenko, Stephanie Pereira, Xianjie Chen, Wenlin Chen, Vijay Rao, Bill Jia, Liang Xiong, and Misha Smelyanskiy · 2019
Cited alongside, same era.
Veloc: Towards high performance adaptive asynchronous checkpointing at large scale
Bogdan Nicolae, Adam Moody, Elsa Gonsiorowski, Kathryn Mohror, and Franck Cappello · 2019
Cited alongside, same era.
PyTorch: An Imperative Style, High-Performance Deep Learning Library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Köpf, Edward Yang, Zach DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala · 2019
Cited alongside, same era.
High-resolution image synthesis with latent diffusion models, 2021
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer · 2021
Later among the works it cites.
Interpreting write performance of supercomputer i/o systems with regression models
Bing Xie, Zilong Tan, Philip Carns, Jeff Chase, Kevin Harms, Jay Lofstead, Sarp Oral, Sudharshan S Vazhkudai, and Feiyi Wang · 2021
Later among the works it cites.
https://pagure.io/libaio , 2016
Linux aio · 2022
Later among the works it cites.
https://www.top500.org/system/179699/ , 2019
Jean zay supercomputer · 2022
Later among the works it cites.
https://github.com/microsoft/DeepSpeed , 2020
Deepspeed · 2022
Later among the works it cites.
https://www.top500.org/system/179842/ , 2020
NVIDIA Selene SuperComputer · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Cited alongside, same era.
Time-based sequence model for personalization and recommendation systems
Tigran Ishkhanov, Maxim Naumov, Xianjie Chen, Yan Zhu, Yuan Zhong, Alisson Gusatti Azzolini, Chonglin Sun, Frank Jiang, Andrey Malevich, and Liang Xiong · 2020
Cited alongside, same era.
Deepfreeze: Towards scalable asynchronous checkpointing of deep learning models
Bogdan Nicolae, Jiali Li, Justin M. Wozniak, George Bosilca, Matthieu Dorier, and Franck Cappello · 2020
Cited alongside, same era.
A study of checkpointing in large scale training of deep neural networks
Elvis Rojas, Albert Njoroge Kahira, Esteban Meneses, Leonardo Bautista Gomez, and Rosa M Badia · 2020
Cited alongside, same era.
Check-in: In-storage checkpointing for key-value store system leveraging flash-based ssds
Joohyeong Yoon, Won Seob Jeong, and Won Woo Ro · 2020
Cited alongside, same era.
Clairvoyant prefetching for distributed machine learning i/o
Nikoli Dryden, Roman Böhringer, Tal Ben-Nun, and Torsten Hoefler · 2021
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2021
Cited alongside, same era.
Gshard: Scaling giant models with conditional computation and automatic sharding
Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen · 2021
Cited alongside, same era.
https://www.weka.io/blog/microsoft-performance-gpudirect/ , 2020
Wekafs · 2022
Later among the works it cites.
Deepspeed-inference: Enabling efficient inference of transformer models at unprecedented scale
R. Aminabadi, S. Rajbhandari, A. Awan, C. Li, D. Li, E. Zheng, O. Ruwase, S. Smith, M. Zhang, J. Rasley, and Y. He · 2022
Later among the works it cites.
Efficient io with io_uring
Jens Axboe · 2022
Later among the works it cites.
Check-N-Run: a checkpointing system for training deep learning recommendation models
Assaf Eisenman, Kiran Kumar Matam, Steven Ingram, Dheevatsa Mudigere, Raghuraman Krishnamoorthi, Krishnakumar Nair, Misha Smelyanskiy, and Murali Annavaram · 2022
Later among the works it cites.
A persistent key-value store for fast storage environments
Meta Inc · 2022
Later among the works it cites.
Hardware/software co-programmable framework for computational ssds to accelerate deep learning service on large-scale graphs
Miryeong Kwon, Donghyun Gouk, Sangwon Lee, and Myoungsoo Jung · 2022
Later among the works it cites.
GPUDirect Storage: A Direct Path Between Storage and GPU Memory
NVIDIA · 2022
Later among the works it cites.
DeepSpeed-MoE: Advancing mixture-of-experts inference and training to power next-generation AI scale
Samyam Rajbhandari, Conglong Li, Zhewei Yao, Minjia Zhang, Reza Yazdani Aminabadi, Ammar Ahmad Awan, Jeff Rasley, and Yuxiong He · 2022
Later among the works it cites.
Shaden Smith, Mostofa Patwary, Brandon Norick, Patrick LeGresley, Samyam Rajbhandari, Jared Casper, Zhun Liu, Shrimai Prabhumoye, George Zerveas, Vijay Korthikanti, Elton Zheng, Rewon Child, Reza Yazdani Aminabadi, Julie Bernauer, Xia Song, Mohammad Shoeybi, Yuxiong He, Michael Houston, Saurabh Tiwary, and Bryan Catanzaro · 2022
Later among the works it cites.
Bloom: A 176b-parameter open-access multilingual language model, 2022
BigScience. Workshop · 2022
Later among the works it cites.