Fetching the paper…
Reading the bibliography…
Lowering costs by driving high utilization across deep learning workloads is a crucial lever for cloud providers.
Process migration in demos/mp
Michael L Powell and Barton P Miller · 1983
Earlier work this paper cites.
Transparent process migration: Design alternatives and the sprite implementation
Fred Douglis and John Ousterhout · 1991
Earlier work this paper cites.
Dmtcp: Transparent checkpointing for cluster computations and the desktop
Jason Ansel, Kapil Arya, and Gene Cooperman · 2009
Earlier work this paper cites.
Optimus: an efficient dynamic resource scheduler for deep learning clusters
Yanghua Peng, Yixin Bao, Yangrui Chen, Chuan Wu, and Chuanxiong Guo · 2018
Earlier work this paper cites.
Mana for mpi: Mpi-agnostic network-agnostic transparent checkpointing
Rohan Garg, Gregory Price, and Gene Cooperman · 2019
Cited alongside, same era.
Zero: Memory optimization towards training A trillion parameter models
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He · 2019
Cited alongside, same era.
Megatron-lm: Training multi-billion parameter language models using gpu model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Cited alongside, same era.
Balancing efficiency and fairness in heterogeneous gpu clusters for deep learning
Shubham Chaudhary, Ramachandran Ramjee, Muthian Sivathanu, Nipun Kwatra, and Srinidhi Viswanatha · 2020
Cited alongside, same era.
https://criu.org/Main_Page
CRIU: Checkpoint Restore in Userspace
Cited in the paper.
https://github.com/microsoft/DeepSpeedExamples/tree/master/Megatron-LM-v1.1.5-3D_parallelism
DeepSpeedExamples/Megatron-LM-3D_parallelism
Cited in the paper.
https://github.com/dnddnjs/pytorch-cifar10/tree/master/pyramidnet
dnddnjs/pytorch-cifar10: pytorch-cifar10/pyramidnet
Cited in the paper.
https://github.com/awslabs/dynamic-training-with-apache-mxnet-on-aws
Dynamic Training with Apache MXNet
Cited in the paper.
https://github.com/huggingface/transformers/tree/master/examples%2Fpytorch%2Ftext-classification
Huggingface/transformers - transformers/examples/pytorch/text-classification’
Cited in the paper.
https://github.com/NVIDIA/apex
Nvidia apex library
Cited in the paper.
In https://docs.nvidia.com/deploy/mps/index.html
Nvidia multi-process service: Gpu management and deployment
Cited in the paper.
In https://developer.nvidia.com/thrust
Nvidia thrust
Cited in the paper.
Resource elasticity in distributed deep learning
Andrew Or, Haoyu Zhang, and Michael Freedman · 2020
Later among the works it cites.
Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He · 2020
Later among the works it cites.
Elastic resource sharing for distributed deep learning
Changho Hwang, Taehyun Kim, Sunghyun Kim, Jinwoo Shin, and KyoungSoo Park · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…