Fetching the paper…
Reading the bibliography…
PHOENIXOS (PHOS) is the first OS service that can concurrently checkpoint and restore (C/R) GPU processes--a fundamental capability for critical tasks such as fault tolerance, process migration, and fast startup.
Keykos architecture
1985
Earlier work this paper cites.
Checkpoint and migration of unix processes in the condor distributed processing system
1997
Earlier work this paper cites.
EROS: a fast capability system
1999
Earlier work this paper cites.
EROS: a fast capability system
1999
Earlier work this paper cites.
A survey of rollback-recovery protocols in message-passing systems
2002
Earlier work this paper cites.
Live migration of virtual machines
2005
Earlier work this paper cites.
Berkeley lab checkpoint/restart (blcr) for linux clusters
2006
Earlier work this paper cites.
Rodinia: A benchmark suite for heterogeneous computing
2009
Earlier work this paper cites.
Linux-cr: Transparent application checkpoint-restart in linux
2010
Earlier work this paper cites.
Scalable smt-based verification of GPU kernel functions
2010
Earlier work this paper cites.
Gpuverify: a verifier for GPU kernels
2012
Earlier work this paper cites.
Verifying GPU kernels by test amplification
2012
Earlier work this paper cites.
GKLEE: concolic verification and test generation for gpus
2012
Earlier work this paper cites.
Parboil: A revised benchmark suite for scientific and commercial throughput computing
2012
Earlier work this paper cites.
Formal analysis of GPU programs with atomics via conflict-directed delay-bounding
2013
Earlier work this paper cites.
A survey of fault tolerance mechanisms and checkpoint/restart implementations for high performance computing systems
2013
Earlier work this paper cites.
The design and implementation of a verification technique for GPU kernels
2015
Earlier work this paper cites.
TVM: an automated end-to-end optimizing compiler for deep learning
2018
Earlier work this paper cites.
Gandiva: Introspective cluster scheduling for deep learning
2018
Earlier work this paper cites.
Cloud programming simplified: A berkeley view on serverless computing
2019
Earlier work this paper cites.
GPU snapshot: checkpoint offloading for gpu-dense systems
2019
Earlier work this paper cites.
Replayable execution optimized for page sharing for a managed runtime environment
2019
Earlier work this paper cites.
PipeSwitch: Fast pipelined context switching for deep learning applications
2020
Earlier work this paper cites.
Catalyzer: Sub-millisecond startup for serverless computing with initialization-less booting
2020
Earlier work this paper cites.
iguard: In-gpu advanced race detection
2021
Earlier work this paper cites.
Checkfreq: Frequent, fine-grained DNN checkpointing
2021
Cited alongside, same era.
The aurora single level store operating system
2021
Cited alongside, same era.
Benchmarking, analysis, and optimization of serverless function snapshots
2021
Cited alongside, same era.
Faasnap: Faas made fast using snapshot-based vms
2022
Cited alongside, same era.
Check-n-run: a checkpointing system for training deep learning recommendation models
2022
Cited alongside, same era.
Singularity: Planet-scale, preemptive and elastic scheduling of AI workloads
2022
Cited alongside, same era.
Creating a communicator
2024
Closest in time.
Cuda toolkit 12.3 downloads
2024
Closest in time.
Cuda toolkit documentation - driver apis
2024
Closest in time.
Implement asynchronous checkpoint saving (with –dist-ckpt-format torch_dist)
2024
Closest in time.
Parallel thread execution isa version 8.4
2024
Closest in time.
Openai gym
2024
Closest in time.
Cuda semantics
2024
Closest in time.
torchvision
2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2022
Cited alongside, same era.
An empirical study on quality issues of deep learning platform
2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
2023
Cited alongside, same era.
Honeycomb: Secure and efficient GPU executions via static validation
2023
Cited alongside, same era.
GEMINI: fast failure recovery in distributed training with in-memory checkpoints
2023
Cited alongside, same era.
No provisioned concurrency: Fast rdma-codesigned remote fork for serverless computing
2023
Cited alongside, same era.
Closest in time.
Bytecheckpoint: A unified checkpointing system for LLM development
2024
Closest in time.
Characterizing network requirements for GPU API remoting in AI applications
2024
Closest in time.
On-demand and parallel checkpoint/restore for gpu applications
2024
Closest in time.
https://docs.nvidia.com/cuda/cuda-binary-utilities/index.html#cu-filt , 2025
cu++filt · 2025
Closest in time.
https://docs.nvidia.com/cuda/gpudirect-rdma/ , 2025
Developing a linux kernel module using gpudirect rdma · 2025
Closest in time.
https://github.com/flashinfer-ai/flashinfer , 2025
FlashInfer: Kernel Library for LLM Serving · 2025
Closest in time.
https://developer.nvidia.com/blog/cuda-graphs/ , 2025
Getting started with cuda graphs · 2025
Closest in time.
Memory changes tracking
2025
Closest in time.
What runs chatgpt inside microsoft’s ai supercomputer
2025
Closest in time.
Basic linear algebra on nvidia gpus
2025
Closest in time.
Context management
2025
Closest in time.
Cuda binary utilities
2025
Closest in time.
Live migration for gpu-accelerated virtual machines
2025
Closest in time.
Nvidia/cuda-checkpoint
2025
Closest in time.
Criugpu: Transparent checkpointing of gpu-accelerated workloads
2025
Closest in time.
Medusa: Accelerating serverless LLM inference with materialization
2025
Closest in time.