Fetching the paper…
Reading the bibliography…
DNN models across many domains continue to grow in size, resulting in high resource requirements for effective training, and unpalatable (and often unaffordable) costs for organizations and research labs across scales.
“Dynamic Mini-batch SGD for Elastic Distributed Training: Learning in the Limbo of Resources”
Haibin Lin et al · 1904
Earlier work this paper cites.
“Well-Read Students Learn Better: The Impact of Student Initialization on Knowledge Distillation”
Iulia Turc, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 1908
Earlier work this paper cites.
“Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism”
Mohammad Shoeybi et al · 1909
Earlier work this paper cites.
“Stochastic Estimation of the Maximum of a Regression Function”
J. Kiefer and J. Wolfowitz · 1952
Earlier work this paper cites.
“A Case for Redundant Arrays of Inexpensive Disks (RAID)”
David. Patterson, Garth Gibson and Randy. Katz · 1988
Earlier work this paper cites.
“Optimization of Collective Communication Operations in MPICH”
Rajeev Thakur, Rolf Rabenseifner and William Gropp · 2005
Earlier work this paper cites.
“A Cost-Aware Elasticity Provisioning System for the Cloud”
Upendra Sharma, Prashant Shenoy, Sambit Sahu and Anees Shaikh · 2011
Earlier work this paper cites.
“Large Scale Distributed Deep Networks”
Jeffrey Dean et al · 2012
Earlier work this paper cites.
“Large Scale Distributed Deep Networks”
Jeffrey Dean et al · 2012
Earlier work this paper cites.
“cuDNN: Efficient Primitives for Deep Learning”, 2014
Sharan Chetlur et al · 2014
Earlier work this paper cites.
“Project Adam: Building an Efficient and Scalable Deep Learning Training System”
Trishul Chilimbi, Yutaka Suzue, Johnson Apacible and Karthik Kalyanaraman · 2014
Earlier work this paper cites.
“Adam: A Method for Stochastic Optimization”, 2014
Diederik. Kingma and Jimmy Ba · 2014
Earlier work this paper cites.
“One weird trick for parallelizing convolutional neural networks”
Alex Krizhevsky · 2014
Earlier work this paper cites.
“Scaling Distributed Machine Learning with the Parameter Server”
Mu Li et al · 2014
Earlier work this paper cites.
“ImageNet Large Scale Visual Recognition Challenge”
Olga Russakovsky et al · 2014
Earlier work this paper cites.
“On parallelizability of stochastic gradient descent for speech DNNS”
F. Seide et al · 2014
Earlier work this paper cites.
“1-Bit Stochastic Gradient Descent and Application to Data-Parallel Distributed Training of Speech DNNs”
Frank Seide et al · 2014
Earlier work this paper cites.
“Resource Elasticity for Large-Scale Machine Learning”
Botong Huang et al · 2015
Earlier work this paper cites.
“SpotCheck: Designing a Derivative IaaS Cloud on the Spot Market”
Prateek Sharma et al · 2015
Earlier work this paper cites.
“Very Deep Convolutional Networks for Large-Scale Image Recognition”
Karen Simonyan and Andrew Zisserman · 2015
Earlier work this paper cites.
“TensorFlow: A System for Large-Scale Machine Learning”
Martı́n Abadi et al · 2016
Earlier work this paper cites.
“Revisiting Distributed Synchronous SGD”
Jianmin Chen, Rajat Monga, Samy Bengio and Rafal Jozefowicz · 2016
Earlier work this paper cites.
“Training Deep Nets with Sublinear Memory Cost”, 2016
Tianqi Chen, Bing Xu, Chiyuan Zhang and Carlos Guestrin · 2016
Earlier work this paper cites.
“GeePS: Scalable Deep Learning on Distributed GPUs with a GPU-Specialized Parameter Server”
Henggang Cui et al · 2016
Earlier work this paper cites.
“Deep Residual Learning for Image Recognition”
Kaiming He, Xiangyu Zhang, Shaoqing Ren and Jian Sun · 2016
Earlier work this paper cites.
“Flint: Batch-Interactive Data-Intensive Processing on Transient Servers”
Prateek Sharma et al · 2016
Cited alongside, same era.
“Google’s Neural Machine Translation System: Bridging the Gap between Human and Machine Translation”
Yonghui Wu et al · 2016
Cited alongside, same era.
“Blending on-demand and spot instances to lower costs for in-memory storage”
Zichen Xu, Christopher Stewart, Nan Deng and Xiaorui Wang · 2016
Cited alongside, same era.
“Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour”
Priya Goyal et al · 2017
Cited alongside, same era.
“Proteus: agile ML elasticity through tiered reliability in dynamic resource markets”
Aaron Harlap et al · 2017
Cited alongside, same era.
“Language Models are Unsupervised Multitask Learners”, 2019
Alec Radford et al · 2019
Later among the works it cites.
“Nexus: A GPU Cluster Engine for Accelerating DNN-Based Video Analysis”
Haichen Shen et al · 2019
Later among the works it cites.
“PipeMare: Asynchronous Pipeline Parallel DNN Training”
Bowen Yang et al · 2019
Later among the works it cites.
“PipeSwitch: Fast Pipelined Context Switching for Deep Learning Applications”
Zhihao Bai, Zhen Zhang, Yibo Zhu and Xin Jin · 2020
Later among the works it cites.
“Lightweight Preemptible Functions”
Sol Boucher, Anuj Kalia, David. Andersen and Michael Kaminsky · 2020
Later among the works it cites.
“Language Models are Few-Shot Learners”
Tom. Brown et al · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“ImageNet Classification with Deep Convolutional Neural Networks”
Alex Krizhevsky, Ilya Sutskever and Geoffrey. Hinton · 2017
Cited alongside, same era.
“Portfolio-Driven Resource Management for Transient Cloud Servers”
Prateek Sharma, David Irwin and Prashant Shenoy · 2017
Cited alongside, same era.
“Poseidon: An Efficient Communication Architecture for Distributed Deep Learning on GPU Clusters”
Hao Zhang et al · 2017
Cited alongside, same era.
“SLAQ: Quality-Driven Scheduling for Distributed Machine Learning”
Haoyu Zhang, Logan Stafman, Andrew Or and Michael. Freedman · 2017
Cited alongside, same era.
“BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2018
Cited alongside, same era.
“Tributary: spot-dancing for elastic services with latency SLOs”
Aaron Harlap et al · 2018
Cited alongside, same era.
“GPipe: Efficient Training of Giant Neural Networks using Pipeline Parallelism”
Yanping Huang et al · 2018
Cited alongside, same era.
“Checkmate: Breaking the Memory Wall with Optimal Tensor Rematerialization”
Paras Jain et al · 2020
Later among the works it cites.
“Themis: Fair and Efficient GPU Cluster Scheduling”
Kshiteej Mahajan et al · 2020
Later among the works it cites.
“Heterogeneity-Aware Cluster Scheduling Policies for Deep Learning Workloads”
Deepak Narayanan et al · 2020
Later among the works it cites.
“Resource Elasticity in Distributed Deep Learning”
Andrew Or, Haoyu Zhang and Michael Freedman · 2020
Later among the works it cites.
“ZeRO: Memory Optimizations toward Training Trillion Parameter Models”
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase and Yuxiong He · 2020
Later among the works it cites.
“Zero: Memory optimizations toward training trillion parameter models”
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase and Yuxiong He · 2020
Later among the works it cites.
“AntMan: Dynamic Scaling on GPU Clusters for Deep Learning”
Wencong Xiao et al · 2020
Later among the works it cites.
“Online Scheduling of Heterogeneous Distributed Machine Learning Jobs”
Qin Zhang et al · 2020
Later among the works it cites.
“Varuna: Scalable, Low-cost Training of Massive Deep Learning Models”
Sanjith Athlur et al · 2021
Later among the works it cites.
“Amazon EC2 Spot Instances Pricing”, https://aws.amazon.com/ec2/spot/pricing/, 2021
AWS · 2021
Later among the works it cites.
“DAPPLE: A Pipelined Data Parallel Approach for Training Large Models”
Shiqing Fan et al · 2021
Later among the works it cites.
“Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity”
William Fedus, Barret Zoph and Noam Shazeer · 2021
Later among the works it cites.
“Elastic Resource Sharing for Distributed Deep Learning”
Changho Hwang et al · 2021
Later among the works it cites.
“Kubernetes: An open-source system for automating deployment, scaling, and management of containerized applications”, https://kubernetes.io/, 2021
2021
Later among the works it cites.
“Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM”
Deepak Narayanan et al · 2021
Later among the works it cites.
“RAS: Continuously Optimized Region-Wide Datacenter Resource Allocation”
Andrew Newell et al · 2021
Later among the works it cites.
“Operating etcd clusters for Kubernetes”, https://kubernetes.io/docs/tasks/administer-cluster/configure-upgrade-etcd/, 2021
2021
Later among the works it cites.
“TorchElastic”, 2021
PyTorch Developers · 2021
Later among the works it cites.
“Dorylus: Affordable, Scalable, and Accurate GNN Training with Distributed CPU Servers and Serverless Threads”
John Thorpe et al · 2021
Later among the works it cites.