Fetching the paper…
Reading the bibliography…
Deep learning training at scale is resource-intensive and time-consuming, often running across hundreds or thousands of GPUs for weeks or months.
Berkeley lab checkpoint/restart (blcr) for linux clusters
Paul H Hargrove and Jason C Duell · 2006
Earlier work this paper cites.
Dmtcp: Transparent checkpointing for cluster computations and the desktop
Jason Ansel, Kapil Arya, and Gene Cooperman · 2009
Earlier work this paper cites.
Checuda: A checkpoint/restart tool for cuda applications
Hiroyuki Takizawa, Katsuto Sato, Kazuhiko Komatsu, and Hiroaki Kobayashi · 2009
Earlier work this paper cites.
Nvcr: A transparent checkpoint-restart library for nvidia cuda
Akira Nukada, Hiroyuki Takizawa, and Satoshi Matsuoka · 2011
Earlier work this paper cites.
Understand Fat Binaries and JIT Caching
Mark Harris · 2013
Earlier work this paper cites.
Mxnet: A flexible and efficient machine learning library for heterogeneous distributed systems, 2015
Tianqi Chen, Mu Li, Yutian Li, Min Lin, Naiyan Wang, Minjie Wang, Tianjun Xiao, Bing Xu, Chiyuan Zhang, and Zheng Zhang · 2015
Earlier work this paper cites.
University of Oxford Advanced Research Computing
Andrew Richards · 2015
Earlier work this paper cites.
TensorFlow: A system for Large-Scale machine learning
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng · 2016
Earlier work this paper cites.
Deep Learning
Ian Goodfellow, Yoshua Bengio, and Aaron Courville · 2016
Earlier work this paper cites.
Tesla P100
NVIDIA · 2017
Earlier work this paper cites.
Serving deep learning models in a serverless platform
Vatche Ishakian, Vinod Muthusamy, and Aleksander Slominski · 2018
Earlier work this paper cites.
Ray: A distributed framework for emerging AI applications
Philipp Moritz, Robert Nishihara, Stephanie Wang, Alexey Tumanov, Richard Liaw, Eric Liang, Melih Elibol, Zongheng Yang, William Paul, Michael I. Jordan, and Ion Stoica · 2018
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
Adam et al. Paszke · 2019
Earlier work this paper cites.
Determinism in deep learning
Duncan Riach · 2019
Earlier work this paper cites.
Analyzing a five-year failure record of a leadership-class supercomputer
Elvis Rojas, Esteban Meneses, Terry Jones, and Don Maxwell · 2019
Earlier work this paper cites.
MArk: Exploiting cloud services for Cost-Effective, SLO-Aware machine learning inference serving
Chengliang Zhang, Minchen Yu, Wei Wang, and Feng Yan · 2019
Earlier work this paper cites.
Batch: machine learning inference serving on serverless platforms with adaptive batching
Ahsan Ali, Riccardo Pinciroli, Feng Yan, and Evgenia Smirni · 2020
Earlier work this paper cites.
Balancing efficiency and fairness in heterogeneous gpu clusters for deep learning
Shubham Chaudhary, Ramachandran Ramjee, Muthian Sivathanu, Nipun Kwatra, and Srinidhi Viswanatha · 2020
Earlier work this paper cites.
Crac: Checkpoint-restart architecture for cuda with streams and uvm
Twinkle Jain and Gene Cooperman · 2020
Earlier work this paper cites.
Deepfreeze: Towards scalable asynchronous checkpointing of deep learning models
Bogdan Nicolae, Jiali Li, Justin M. Wozniak, George Bosilca, Matthieu Dorier, and Franck Cappello · 2020
Earlier work this paper cites.
A100 Tensor Core GPU Architecture
NVIDIA · 2020
Earlier work this paper cites.
Zero: Memory optimizations toward training trillion parameter models
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He · 2020
Cited alongside, same era.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, et al · 2021
Cited alongside, same era.
CheckFreq: Frequent, Fine-Grained DNN checkpointing
Jayashree Mohan, Amar Phanishayee, and Vijay Chidambaram · 2021
Cited alongside, same era.
drm/amdkfd: CRIU Introduce Checkpoint-Restore APIs, 2021
Rajneesh Bhardwaj · 2021
Cited alongside, same era.
INFaaS: Automated model-less inference serving
Francisco Romero, Qian Li, Neeraja J. Yadwadkar, and Christos Kozyrakis · 2021
Cited alongside, same era.
Serving heterogeneous machine learning models on Multi-GPU servers with Spatio-Temporal sharing
Introduction to LLM Agents
NVIDIA · 2023
Later among the works it cites.
Asyfunc: A high-performance and resource-efficient serverless inference system via asymmetric functions
Qiangyu Pei, Yongjie Yuan, Haichuan Hu, Qiong Chen, and Fangming Liu · 2023
Later among the works it cites.
Fast checkpoint restore for gpus
David Yat Sin Rajneesh Bhardwaj, Felix Kuehling · 2023
Later among the works it cites.
Gemini: Fast failure recovery in distributed training with in-memory checkpoints
Zhuang Wang, Zhen Jia, Shuai Zheng, Zhen Zhang, Xinwei Fu, T. S. Eugene Ng, and Yida Wang · 2023
Later among the works it cites.
ROCm Examples
AMD · 2024
Later among the works it cites.
CRIU Statistics
CRIU · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Seungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park, Youngjin Kwon, and Jaehyuk Huh · 2022
Cited alongside, same era.
Cricket: A virtualization layer for distributed execution of cuda applications with checkpoint/restart support
Niklas Eiling, Jonas Baude, Stefan Lankes, and Antonello Monti · 2022
Cited alongside, same era.
Check-N-Run: a checkpointing system for training deep learning recommendation models
Assaf Eisenman, Kiran Kumar Matam, Steven Ingram, Dheevatsa Mudigere, Raghuraman Krishnamoorthi, Krishnakumar Nair, Misha Smelyanskiy, and Murali Annavaram · 2022
Cited alongside, same era.
Deep reinforcement learning for autonomous driving: A survey
B Ravi Kiran, Ibrahim Sobh, Victor Talpaert, Patrick Mannion, Ahmad A. Al Sallab, Senthil Yogamani, and Patrick Pérez · 2022
Cited alongside, same era.
Tetris: Memory-efficient serverless inference through tensor sharing
Jie Li, Laiping Zhao, Yanan Yang, Kunlin Zhan, and Keqiu Li · 2022
Cited alongside, same era.
H100 Tensor Core GPU Architecture
NVIDIA · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, et al · 2022
Cited alongside, same era.
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, et al · 2024
Later among the works it cites.
Just-in-time checkpointing: Low cost error recovery from deep learning training failures
Tanmaey Gupta, Sanjeev Krishnan, Rituraj Kumar, Abhishek Vijeev, Bhargav Gulavani, Nipun Kwatra, Ramachandran Ramjee, and Muthian Sivathanu · 2024
Later among the works it cites.
Checkpointing cuda applications with criu
Steven Gurfinkel · 2024
Later among the works it cites.
Boost Checkpoint Speed and Reduce Cost with Nebula
Microsoft · 2024
Later among the works it cites.
Driver API Documentation
NVIDIA · 2024
Later among the works it cites.
Megatron-lm
NVIDIA · 2024
Later among the works it cites.
Runtime API Documentation
NVIDIA · 2024
Later among the works it cites.
Code llama: Open foundation models for code, 2024
Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Romain Sauvestre, Tal Remez, Jérémy Rapin, Artyom Kozhevnikov, Ivan Evtimov, Joanna Bitton, Manish Bhatt, Cristian Canton Ferrer, Aaron Grattafiori, Wenhan Xiong, Alexandre Défossez, Jade Copet, Faisal Azhar, Hugo Touvron, Louis Martin, Nicolas Usunier, Thomas Scialom, and Gabriel Synnaeve · 2024
Later among the works it cites.
Gemini: A Family of Highly Capable Multimodal Models, 2024
Gemini Team · 2024
Later among the works it cites.
On-demand and parallel checkpoint/restore for gpu applications
Yanning Yang, Dong Du, Haitao Song, and Yubin Xia · 2024
Later among the works it cites.
Part-time power measurements: nvidia-smi’s lack of attention, 2024
Zeyu Yang, Karel Adamek, and Wesley Armour · 2024
Later among the works it cites.
Memory snapshots: Checkpoint/restore for sub-second startup
Jonathon Belotti · 2025
Closest in time.
ptrace(2) - Linux manual page
The Linux Foundation · 2025
Closest in time.
Github copilot
Microsoft · 2025
Closest in time.
Memverge memory machine ai gpu-as-a-service
Bernie Wu · 2025
Closest in time.