Fetching the paper…
Reading the bibliography…
Efficient GPU resource scheduling is essential to maximize resource utilization and save training costs for the increasing amount of deep learning workloads in shared GPU clusters.
Individual comparisons by ranking methods
Frank Wilcoxon. 1992 · 1992
Earlier work this paper cites.
Connection scheduling in web servers
Mark E Crovella, Robert Frangioso, and Mor Harchol-Balter. 1999 · 1999
Earlier work this paper cites.
CIFAR-10 and CIFAR-100 datasets
Alex Krizhevsky. 2009 · 2009
Earlier work this paper cites.
Performance evaluation of bag of gangs scheduling in a heterogeneous distributed system
Zafeirios C Papazachos and Helen D Karatza. 2010 · 2010
Earlier work this paper cites.
Large scale distributed deep networks. In Advances in neural information processing systems . 1223–1231
Jeffrey Dean, Greg Corrado, Rajat Monga, Kai Chen, Matthieu Devin, Mark Mao, Marc’aurelio Ranzato, Andrew Senior, Paul Tucker, Ke Yang, et al · 2012
Earlier work this paper cites.
Multi-Resource Packing for Cluster Schedulers
Robert Grandl, Ganesh Ananthanarayanan, Srikanth Kandula, Sriram Rao, and Aditya Akella. 2014 · 2014
Earlier work this paper cites.
Very deep convolutional networks for large-scale image recognition
Karen Simonyan and Andrew Zisserman. 2014 · 2014
Earlier work this paper cites.
Going deeper with convolutions. In Proceedings of the IEEE conference on computer vision and pattern recognition . 1–9
Christian Szegedy, Wei Liu, Yangqing Jia, Pierre Sermanet, Scott Reed, Dragomir Anguelov, Dumitru Erhan, Vincent Vanhoucke, and Andrew Rabinovich. 2015 · 2015
Earlier work this paper cites.
Large-scale cluster management at Google with Borg. In Proceedings of the Tenth European Conference on Computer Systems . 1–17
Abhishek Verma, Luis Pedrosa, Madhukar Korupolu, David Oppenheimer, Eric Tune, and John Wilkes. 2015 · 2015
Earlier work this paper cites.
Tensorflow: A system for large-scale machine learning. In 12th { \{ USENIX } \} symposium on operating systems design and implementation ( { \{ OSDI } \} 16) . 265–283
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al · 2016
Earlier work this paper cites.
Deep speech 2: End-to-end speech recognition in english and mandarin. In International conference on machine learning . 173–182
Dario Amodei, Sundaram Ananthanarayanan, Rishita Anubhai, Jingliang Bai, Eric Battenberg, Carl Case, Jared Casper, Bryan Catanzaro, Qiang Cheng, Guoliang Chen, et al · 2016
Earlier work this paper cites.
Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition . 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
On large-batch training for deep learning: Generalization gap and sharp minima
Nitish Shirish Keskar, Dheevatsa Mudigere, Jorge Nocedal, Mikhail Smelyanskiy, and Ping Tak Peter Tang. 2016 · 2016
Earlier work this paper cites.
ImageNet
Stanford Vision Lab. 2016 · 2016
Cited alongside, same era.
Resource management with deep reinforcement learning. In Proceedings of the 15th ACM Workshop on Hot Topics in Networks . ACM, 50–56
Hongzi Mao, Mohammad Alizadeh, Ishai Menache, and Srikanth Kandula. 2016 · 2016
Cited alongside, same era.
Rethinking the inception architecture for computer vision. In Proceedings of the IEEE conference on computer vision and pattern recognition . 2818–2826
Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. 2016 · 2016
Cited alongside, same era.
Accurate, large minibatch sgd: Training imagenet in 1 hour
Priya Goyal, Piotr Dollár, Ross Girshick, Pieter Noordhuis, Lukasz Wesolowski, Aapo Kyrola, Andrew Tulloch, Yangqing Jia, and Kaiming He. 2017 · 2017
Cited alongside, same era.
Train longer, generalize better: closing the generalization gap in large batch training of neural networks. In Advances in Neural Information Processing Systems . 1731–1741
Gandiva: Introspective cluster scheduling for deep learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18) . 595–610
Wencong Xiao, Romil Bhardwaj, Ramachandran Ramjee, Muthian Sivathanu, Nipun Kwatra, Zhenhua Han, Pratyush Patel, Xuan Peng, Hanyu Zhao, Quanlu Zhang, et al · 2018
Later among the works it cites.
Yang You, Zhao Zhang, Cho-Jui Hsieh, James Demmel, and Kurt Keutzer. 2018 · 2018
Later among the works it cites.
Chic: experience-driven scheduling in machine learning clusters. In Proceedings of the International Symposium on Quality of Service . 1–10
Yifan Gong, Baochun Li, Ben Liang, and Zheng Zhan. 2019 · 2019
Later among the works it cites.
Tiresias: A GPU Cluster Manager for Distributed Deep Learning. In 16th USENIX Symposium on Networked Systems Design and Implementation (NSDI 19) . 485–500
Juncheng Gu, Mosharaf Chowdhury, Kang G Shin, Yibo Zhu, Myeongjae Jeon, Junjie Qian, Hongqiang Liu, and Chuanxiong Guo. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Elad Hoffer, Itay Hubara, and Daniel Soudry. 2017 · 2017
Cited alongside, same era.
Cyclical learning rates for training neural networks. In 2017 IEEE Winter Conference on Applications of Computer Vision (WACV) . IEEE, 464–472
Leslie N Smith. 2017 · 2017
Cited alongside, same era.
Don’t decay the learning rate, increase the batch size
Samuel L Smith, Pieter-Jan Kindermans, Chris Ying, and Quoc V Le. 2017 · 2017
Cited alongside, same era.
Large batch training of convolutional networks
Yang You, Igor Gitman, and Boris Ginsburg. 2017 · 2017
Cited alongside, same era.
Haoyu Zhang, Logan Stafman, Andrew Or, and Michael J Freedman. 2017 · 2017
Cited alongside, same era.
Online Job Scheduling in Distributed Machine Learning Clusters. In IEEE INFOCOM 2018 - IEEE Conference on Computer Communications . 495–503
Y. Bao, Y. Peng, C. Wu, and Z. Li. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Cited alongside, same era.
SRPT for multiserver systems
Isaac Grosof, Ziv Scully, and Mor Harchol-Balter. 2018 · 2018
Cited alongside, same era.
Analysis of Large-Scale Multi-Tenant GPU Clusters for DNN Training Workloads. In 2019 USENIX Annual Technical Conference (USENIXATC 19)
Myeongjae Jeon, Shivaram Venkataraman, Amar Phanishayee, Junjie Qian, Wencong Xiao, and Fan Yang. 2019 · 2019
Later among the works it cites.
DRAGON: A Dynamic Scheduling and Scaling Controller for Managing Distributed Deep Learning Jobs in Kubernetes Cluster. In Proceedings of the 9th International Conference on Cloud Computing and Services Science, CLOSER 2019, Heraklion, Crete, Greece, May 2-4, 2019 . 569–577
Chan-Yi Lin, Ting-An Yeh, and Jerry Chou. 2019a · 2019
Later among the works it cites.
Dynamic mini-batch SGD for elastic distributed training: learning in the limbo of resources
Haibin Lin, Hang Zhang, Yifei Ma, Tong He, Zhi Zhang, Sheng Zha, and Mu Li. 2019b · 2019
Later among the works it cites.
Learning scheduling algorithms for data processing clusters
Hongzi Mao, Malte Schwarzkopf, Shaileshh Bojja Venkatakrishnan, Zili Meng, and Mohammad Alizadeh. 2019 · 2019
Later among the works it cites.
DL2: A Deep Learning-driven Scheduler for Deep Learning Clusters
Yanghua Peng, Yixin Bao, Yangrui Chen, Chuan Wu, Chen Meng, and Wei Lin. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
Large batch optimization for deep learning: Training bert in 76 minutes
Yang You, Jing Li, Sashank Reddi, Jonathan Hseu, Sanjiv Kumar, Srinadh Bhojanapalli, Xiaodan Song, James Demmel, Kurt Keutzer, and Cho-Jui Hsieh. 2019 · 2019
Later among the works it cites.
NVIDIA Collective Communications Library (NCCL)
NVIDIA. 2020 · 2020
Later among the works it cites.
LONGHORN - TEXAS ADVANCED COMPUTING CENTER
Texas Advanced Computing Center. 2021 · 2021
Closest in time.