Fetching the paper…
Reading the bibliography…
Deep learning (DL) shows its prosperity in a wide variety of fields.
An experimental time-sharing system. In spring joint computer conference
Fernando J Corbató, Marjorie Merwin-Daggett, and Robert C Daley. 1962 · 1962
Earlier work this paper cites.
An analysis of decay-usage scheduling in multiprocessors
Dick HJ Epema. 1995 · 1995
Earlier work this paper cites.
Parallel job scheduling: Issues and approaches. In Workshop on Job Scheduling Strategies for Parallel Processing
Dror G Feitelson and Larry Rudolph. 1995 · 1995
Earlier work this paper cites.
Packing schemes for gang scheduling. In Job Scheduling Strategies for Parallel Processing
Dror G. Feitelson. 1996 · 1996
Earlier work this paper cites.
Job scheduling in multiprogrammed parallel systems
Dror G Feitelson. 1997 · 1997
Earlier work this paper cites.
Theory and practice in parallel job scheduling. In Workshop on Job Scheduling Strategies for Parallel Processing
Dror G Feitelson, Larry Rudolph, Uwe Schwiegelshohn, Kenneth C Sevcik, and Parkson Wong. 1997 · 1997
Earlier work this paper cites.
Utilization, predictability, workloads, and user runtime estimates in scheduling the IBM SP2 with backfilling
A.W. Mu’alem and D.G. Feitelson. 2001 · 2001
Earlier work this paper cites.
SLURM: Simple Linux Utility for Resource Management. In Job Scheduling Strategies for Parallel Processing
Andy B. Yoo, Morris A. Jette, and Mark Grondona. 2003 · 2003
Earlier work this paper cites.
Delay scheduling: a simple technique for achieving locality and fairness in cluster scheduling. In Proceedings of the 5th European conference on Computer systems . 265–278
Matei Zaharia, Dhruba Borthakur, Joydeep Sen Sarma, Khaled Elmeleegy, Scott Shenker, and Ion Stoica. 2010 · 2010
Earlier work this paper cites.
Dominant Resource Fairness: Fair Allocation of Multiple Resource Types. In Proceedings of the 8th USENIX Conference on Networked Systems Design and Implementation (NSDI ’11)
Ali Ghodsi, Matei Zaharia, Benjamin Hindman, Andy Konwinski, Scott Shenker, and Ion Stoica. 2011 · 2011
Earlier work this paper cites.
Mesos: A Platform for Fine-Grained Resource Sharing in the Data Center. In 8th USENIX Symposium on Networked Systems Design and Implementation (NSDI ’11)
Benjamin Hindman, Andy Konwinski, Matei Zaharia, Ali Ghodsi, Anthony D. Joseph, Randy Katz, Scott Shenker, and Ion Stoica. 2011 · 2011
Earlier work this paper cites.
vCUDA: GPU-Accelerated High-Performance Computing in Virtual Machines
Lin Shi, Hao Chen, Jianhua Sun, and Kenli Li. 2012 · 2012
Earlier work this paper cites.
Apache Hadoop YARN: Yet Another Resource Negotiator. In Proceedings of the 4th Annual Symposium on Cloud Computing (SoCC ’13)
Vinod Kumar Vavilapalli, Arun C. Murthy, Chris Douglas, Sharad Agarwal, Mahadev Konar, Robert Evans, Thomas Graves, Jason Lowe, Hitesh Shah, Siddharth Seth, Bikas Saha, Carlo Curino, Owen O’Malley, Sanjay Radia, Benjamin Reed, and Eric Baldeschwieler. 2013 · 2013
Earlier work this paper cites.
Efficient coflow scheduling without prior knowledge
Mosharaf Chowdhury and Ion Stoica. 2015 · 2015
Earlier work this paper cites.
Cloud Computing Resource Scheduling and a Survey of Its Evolutionary Approaches
Zhi-Hui Zhan, Xiao-Fang Liu, Yue-Jiao Gong, Jun Zhang, Henry Shu-Hung Chung, and Yun Li. 2015 · 2015
Earlier work this paper cites.
TensorFlow: A System for Large-Scale Machine Learning. In 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI ’16)
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. 2016 · 2016
Earlier work this paper cites.
Borg, Omega, and Kubernetes: Lessons Learned from Three Container-Management Systems over a Decade
Brendan Burns, Brian Grant, David Oppenheimer, Eric Brewer, and John Wilkes. 2016 · 2016
Earlier work this paper cites.
SERF: Efficient Scheduling for Fast Deep Neural Network Serving via Judicious Parallelism. In SC ’16: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis
Feng Yan, Olatunji Ruwase, Yuxiong He, and Evgenia Smirni. 2016 · 2016
Earlier work this paper cites.
Topology-Aware GPU Scheduling for Learning Workloads in Cloud Environments. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC ’17)
Marcelo Amaral, Jordà Polo, David Carrera, Seetharami Seelam, and Malgorzata Steinder. 2017 · 2017
Earlier work this paper cites.
Clipper: A Low-Latency Online Prediction Serving System. In 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17)
Daniel Crankshaw, Xin Wang, Guilio Zhou, Michael J. Franklin, Joseph E. Gonzalez, and Ion Stoica. 2017 · 2017
Earlier work this paper cites.
Model-agnostic meta-learning for fast adaptation of deep networks. In International conference on machine learning
Chelsea Finn, Pieter Abbeel, and Sergey Levine. 2017 · 2017
Earlier work this paper cites.
Deep reinforcement learning for robotic manipulation with asynchronous off-policy updates. In 2017 IEEE international conference on robotics and automation (ICRA) . IEEE, 3389–3396
Shixiang Gu, Ethan Holly, Timothy Lillicrap, and Sergey Levine. 2017 · 2017
Earlier work this paper cites.
Swayam: distributed autoscaling to meet SLAs of machine learning inference services with resource efficiency. In Proceedings of the 18th ACM/IFIP/USENIX Middleware Conference (Middleware ’17)
Arpan Gujarati, Sameh Elnikety, Yuxiong He, Kathryn S. McKinley, and Björn B. Brandenburg. 2017 · 2017
Earlier work this paper cites.
GPU Virtualization and Scheduling Methods: A Comprehensive Survey
Cheol-Ho Hong, Ivor Spence, and Dimitrios S. Nikolopoulos. 2017 · 2017
Earlier work this paper cites.
TensorFlow-Serving: Flexible, High-Performance ML Serving
Christopher Olston, Noah Fiedel, Kiril Gorovoy, Jeremiah Harmsen, Li Lao, Fangwei Li, Vinu Rajashekhar, Sukriti Ramesh, and Jordan Soyke. 2017 · 2017
Earlier work this paper cites.
HyperDrive: Exploring Hyperparameters with POP Scheduling. In Proceedings of the 18th International Middleware Conference (Middleware ’17)
Jeff Rasley, Yuxiong He, Feng Yan, Olatunji Ruwase, and Rodrigo Fonseca. 2017 · 2017
Earlier work this paper cites.
Deep learning in medical image analysis
Dinggang Shen, Guorong Wu, and Heung-Il Suk. 2017 · 2017
Earlier work this paper cites.
Workload scheduling in cloud: A comprehensive survey and future research directions. In International Conference on Cloud Computing, Data Science Engineering - Confluence
S R Shishira, A. Kandasamy, and K. Chandrasekaran. 2017 · 2017
Earlier work this paper cites.
Mastering the game of go without human knowledge
David Silver, Julian Schrittwieser, Karen Simonyan, Ioannis Antonoglou, Aja Huang, Arthur Guez, Thomas Hubert, Lucas Baker, Matthew Lai, Adrian Bolton, et al · 2017
Earlier work this paper cites.
Towards distributed machine learning in shared clusters: A dynamically-partitioned approach. In 2017 IEEE International Conference on Smart Computing (SMARTCOMP ’17)
Peng Sun, Yonggang Wen, Nguyen Binh Duong Ta, and Shengen Yan. 2017 · 2017
Earlier work this paper cites.
Operating systems: Three easy pieces
Remzi H Arpaci-Dusseau and Andrea C Arpaci-Dusseau. 2018 · 2018
Earlier work this paper cites.
Online job scheduling in distributed machine learning clusters. In IEEE INFOCOM 2018-IEEE Conference on Computer Communications (INFOCOM ’18)
Yixin Bao, Yanghua Peng, Chuan Wu, and Zongpeng Li. 2018 · 2018
Earlier work this paper cites.
Low Latency RNN Inference with Cellular Batching. In Proceedings of the Thirteenth EuroSys Conference (EuroSys ’18)
Pin Gao, Lingfan Yu, Yongwei Wu, and Jinyang Li. 2018 · 2018
Earlier work this paper cites.
GPGPU Power Modeling for Multi-domain Voltage-Frequency Scaling. In 2018 IEEE International Symposium on High Performance Computer Architecture (HPCA ’18)
Joao Guerreiro, Aleksandar Ilic, Nuno Roma, and Pedro Tomas. 2018 · 2018
Earlier work this paper cites.
Applied machine learning at facebook: A datacenter infrastructure perspective. In 2018 IEEE International Symposium on High Performance Computer Architecture (HPCA ’18)
Kim Hazelwood, Sarah Bird, David Brooks, Soumith Chintala, Utku Diril, Dmytro Dzhulgakov, Mohamed Fawzy, Bill Jia, Yangqing Jia, Aditya Kalro, et al · 2018
Earlier work this paper cites.
Serving Deep Learning Models in a Serverless Platform. In 2018 IEEE International Conference on Cloud Engineering (IC2E)
Vatche Ishakian, Vinod Muthusamy, and Aleksander Slominski. 2018 · 2018
Earlier work this paper cites.
Dynamic space-time scheduling for gpu inference
Paras Jain, Xiangxi Mo, Ajay Jain, Harikaran Subbaraj, Rehan Sohail Durrani, Alexey Tumanov, Joseph Gonzalez, and Ion Stoica. 2018 · 2018
Earlier work this paper cites.
PRETZEL: Opening the Black Box of Machine Learning Prediction Serving Systems. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI ’18)
Yunseong Lee, Alberto Scolari, Byung-Gon Chun, Marco Domenico Santambrogio, Markus Weimer, and Matteo Interlandi. 2018 · 2018
Earlier work this paper cites.
Ease.Ml: Towards Multi-Tenant Resource Sharing for Machine Learning Workloads
Tian Li, Jie Zhong, Ji Liu, Wentao Wu, and Ce Zhang. 2018 · 2018
Earlier work this paper cites.
HPC Cloud for Scientific and Business Applications: Taxonomy, Vision, and Research Challenges
Marco A. S. Netto, Rodrigo N. Calheiros, Eduardo R. Rodrigues, Renato L. F. Cunha, and Rajkumar Buyya. 2018 · 2018
Earlier work this paper cites.
Deep Learning Inference in Facebook Data Centers: Characterization, Performance Optimizations and Hardware Implications
Jongsoo Park, Maxim Naumov, Protonu Basu, Summer Deng, Aravind Kalaiah, Daya Khudia, James Law, Parth Malani, Andrey Malevich, Satish Nadathur, Juan Pino, Martin Schatz, Alexander Sidorov, Viswanath Sivakumar, Andrew Tulloch, Xiaodong Wang, Yiming Wu, Hector Yuen, Utku Diril, Dmytro Dzhulgakov, Kim Hazelwood, Bill Jia, Yangqing Jia, Lin Qiao, Vijay Rao, Nadav Rotem, Sungjoo Yoo, and Mikhail Smelyanskiy. 2018 · 2018
Earlier work this paper cites.
Optimus: An Efficient Dynamic Resource Scheduler for Deep Learning Clusters. In Proceedings of the Thirteenth EuroSys Conference (EuroSys ’18)
Yanghua Peng, Yixin Bao, Yangrui Chen, Chuan Wu, and Chuanxiong Guo. 2018 · 2018
Earlier work this paper cites.
Scalable system scheduling for HPC and big data
Albert Reuther, Chansup Byun, William Arcand, David Bestor, Bill Bergeron, Matthew Hubbell, Michael Jones, Peter Michaleas, Andrew Prout, Antonio Rosa, and Jeremy Kepner. 2018 · 2018
Earlier work this paper cites.
Horovod: fast and easy distributed deep learning in TensorFlow
Alexander Sergeev and Mike Del Balso. 2018 · 2018
Earlier work this paper cites.
Rafiki: Machine Learning as an Analytics Service System
Wei Wang, Jinyang Gao, Meihui Zhang, Sheng Wang, Gang Chen, Teck Khim Ng, Beng Chin Ooi, Jie Shao, and Moaz Reyad. 2018 · 2018
Earlier work this paper cites.
Gandiva: Introspective Cluster Scheduling for Deep Learning. In USENIX Symposium on Operating Systems Design and Implementation
Wencong Xiao, Romil Bhardwaj, Ramachandran Ramjee, Muthian Sivathanu, Nipun Kwatra, Zhenhua Han, Pratyush Patel, Xuan Peng, Hanyu Zhao, Quanlu Zhang, Fan Yang, and Lidong Zhou. 2018 · 2018
Earlier work this paper cites.
Poseidon: An efficient communication architecture for distributed deep learning on GPU clusters. In 2018 USENIX Annual Technical Conference (ATC ’18)
Hao Zhang, Zeyu Zheng, Shizhen Xu, Wei Dai, Qirong Ho, Xiaodan Liang, Zhiting Hu, Jinliang Wei, Pengtao Xie, and Eric P Xing. 2017 · 2018
Earlier work this paper cites.
Deep Learning-based Job Placement in Distributed Machine Learning Clusters. In IEEE Conference on Computer Communications (INFOCOM ’19)
Yixin Bao, Yanghua Peng, and Chuan Wu. 2019 · 2019
Earlier work this paper cites.
BARISTA: Efficient and Scalable Serverless Serving System for Deep Learning Prediction Services. In 2019 IEEE International Conference on Cloud Engineering (IC2E)
Anirban Bhattacharjee, Ajay Dev Chhokra, Zhuangwei Kang, Hongyang Sun, Aniruddha Gokhale, and Gabor Karsai. 2019 · 2019
Earlier work this paper cites.
Ebird: Elastic Batch for Improving Responsiveness and Throughput of Deep Learning Services. In 2019 IEEE 37th International Conference on Computer Design (ICCD)
Weihao Cui, Mengze Wei, Quan Chen, Xiaoxin Tang, Jingwen Leng, Li Li, and Mingyi Guo. 2019 · 2019
Earlier work this paper cites.
TrIMS: Transparent and Isolated Model Sharing for Low Latency Deep Learning Inference in Function-as-a-Service. In 2019 IEEE 12th International Conference on Cloud Computing (CLOUD)
Abdul Dakkak, Cheng Li, Simon Garcia de Gonzalo, Jinjun Xiong, and Wen-mei Hwu. 2019 · 2019
Earlier work this paper cites.
Challenges and Opportunities of DNN Model Execution Caching. In Proceedings of the Workshop on Distributed Infrastructures for Deep Learning (DIDL ’19)
Guin R. Gilman, Samuel S. Ogden, Robert J. Walls, and Tian Guo. 2019 · 2019
Earlier work this paper cites.
Tiresias: A GPU Cluster Manager for Distributed Deep Learning. In 16th USENIX Symposium on Networked Systems Design and Implementation (NSDI ’19)
Juncheng Gu, Mosharaf Chowdhury, Kang G. Shin, Yibo Zhu, Myeongjae Jeon, Junjie Qian, Hongqiang Liu, and Chuanxiong Guo. 2019 · 2019
Earlier work this paper cites.
One Size Does Not Fit All: Quantifying and Exposing the Accuracy-Latency Trade-Off in Machine Learning Cloud Service APIs via Tolerance Tiers. In 2019 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS)
Matthew Halpern, Behzad Boroujerdian, Todd Mummert, Evelyn Duesterwald, and Vijay Janapa Reddi. 2019 · 2019
Cited alongside, same era.
GRNN: Low-Latency and Scalable RNN Inference on GPUs. In Proceedings of the Fourteenth EuroSys Conference 2019 (EuroSys ’19)
Connor Holmes, Daniel Mawhirter, Yuxiong He, Feng Yan, and Bo Wu. 2019 · 2019
Cited alongside, same era.
Deliver high performance ML inference with AWS Inferentia
Gadi Hutt, Vibhav Viswanathan, and Adam Nadolski. 2019 · 2019
Cited alongside, same era.
Optimizing on-demand gpus in the cloud for deep learning applications training. In 2019 4th International Conference on Computing, Communications and Security (ICCCS ’18)
Arezoo Jahani, Marco Lattuada, Michele Ciavotta, Danilo Ardagna, Edoardo Amaldi, and Li Zhang. 2019 · 2019
Cited alongside, same era.
Enabling Cost-Effective, SLO-Aware Machine Learning Inference Serving on Public Cloud
Chengliang Zhang, Minchen Yu, Feng Yan, et al · 2020
Later among the works it cites.
DyBatch: Efficient Batching and Fair Scheduling for Deep Learning Inference on Time-sharing Devices. In 2020 20th IEEE/ACM International Symposium on Cluster, Cloud and Internet Computing (CCGRID)
Shaojun Zhang, Wei Li, Chen Wang, Zahir Tari, and Albert Y. Zomaya. 2020c · 2020
Later among the works it cites.
CODA: Improving Resource Utilization by Slimming and Co-locating DNN and CPU Jobs. In 2020 IEEE 40th International Conference on Distributed Computing Systems (ICDCS ’20)
Han Zhao, Weihao Cui, Quan Chen, Jingwen Leng, Kai Yu, Deze Zeng, Chao Li, and Minyi Guo. 2020a · 2020
Later among the works it cites.
JPAS: Job-progress-aware flow scheduling for deep learning clusters
Pan Zhou, Xinshu He, Shouxi Luo, Hongfang Yu, and Gang Sun. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
FfDL: A Flexible Multi-tenant Deep Learning Platform. In Proceedings of the 20th International Middleware Conference (Middleware ’19)
K. R. Jayaram, Vinod Muthusamy, Parijat Dube, Vatche Ishakian, Chen Wang, Benjamin Herta, Scott Boag, Diana Arroyo, Asser Tantawi, Archit Verma, Falk Pollok, and Rania Khalaf. 2019 · 2019
Cited alongside, same era.
Analysis of Large-Scale Multi-Tenant GPU Clusters for DNN Training Workloads. In 2019 USENIX Annual Technical Conference (USENIX ATC ’19)
Myeongjae Jeon, Shivaram Venkataraman, Amar Phanishayee, Junjie Qian, Wencong Xiao, and Fan Yang. 2019 · 2019
Cited alongside, same era.
Privacy Accounting and Quality Control in the Sage Differentially Private ML Platform. In Proceedings of the 27th ACM Symposium on Operating Systems Principles (SOSP ’19)
Mathias Lécuyer, Riley Spahn, Kiran Vodrahalli, Roxana Geambasu, and Daniel Hsu. 2019 · 2019
Cited alongside, same era.
HyperSched: Dynamic Resource Reallocation for Model Development on a Deadline. In Proceedings of the ACM Symposium on Cloud Computing (SoCC ’19)
Richard Liaw, Romil Bhardwaj, Lisa Dunlap, Yitian Zou, Joseph E. Gonzalez, Ion Stoica, and Alexey Tumanov. 2019 · 2019
Cited alongside, same era.
DRAGON: A Dynamic Scheduling and Scaling Controller for Managing Distributed Deep Learning Jobs in Kubernetes Cluster. In CLOSER (CLOSER ’19)
Chan-Yi Lin, Ting-An Yeh, and Jerry Chou. 2019 · 2019
Cited alongside, same era.
SCHED 2 : Scheduling Deep Learning Training via Deep Reinforcement Learning. In 2019 IEEE Global Communications Conference (GLOBECOM ’19)
Yunteng Luan, Xukun Chen, Hanyu Zhao, Zhi Yang, and Yafei Dai. 2019 · 2019
Cited alongside, same era.
PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Advances in Neural Information Processing Systems
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019 · 2019
Cited alongside, same era.
Swift machine learning model serving scheduling: a region based reinforcement learning approach. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC ’19)
Heyang Qin, Syed Zawad, Yanqi Zhou, Lei Yang, Dongfang Zhao, and Feng Yan. 2019 · 2019
Cited alongside, same era.
Online evolutionary batch size orchestration for scheduling deep learning workloads in GPU clusters. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC ’21)
Zhengda Bian, Shenggui Li, Wei Wang, and Yang You. 2021 · 2021
Later among the works it cites.
Switches for HIRE: Resource Scheduling for Data Center in-Network Computing. In Proceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS ’21)
Marcel Blöcher, Lin Wang, Patrick Eugster, and Max Schmidt. 2021 · 2021
Later among the works it cites.
Performance and Cost Comparison of Cloud Services for Deep Learning Workload. In Companion of the ACM/SPEC International Conference on Performance Engineering
Dheeraj Chahal, Mayank Mishra, Surya Palepu, and Rekha Singhal. 2021 · 2021
Later among the works it cites.
DynamoML: Dynamic Resource Management Operators for Machine Learning Workloads.. In CLOSER (CLOSER ’21)
Min-Chi Chiang and Jerry Chou. 2021 · 2021
Later among the works it cites.
E2bird: Enhanced Elastic Batch for Improving Responsiveness and Throughput of Deep Learning Services
Weihao Cui, Quan Chen, Han Zhao, Mengze Wei, Xiaoxin Tang, and Minyi Guo. 2021a · 2021
Later among the works it cites.
IOS: Inter-Operator Scheduler for CNN Acceleration. In Proceedings of Machine Learning and Systems (MLSys ’21)
Yaoyao Ding, Ligeng Zhu, Zhihao Jia, Gennady Pekhimenko, and Song Han. 2021 · 2021
Later among the works it cites.
Elastic Hyperparameter Tuning on the Cloud. In Proceedings of the ACM Symposium on Cloud Computing (SoCC ’21)
Lisa Dunlap, Kirthevasan Kandasamy, Ujval Misra, Richard Liaw, Michael Jordan, Ion Stoica, and Joseph E. Gonzalez. 2021 · 2021
Later among the works it cites.
ANDREAS: Artificial intelligence traiNing scheDuler foR accElerAted resource clusterS. In 2021 8th International Conference on Future Internet of Things and Cloud (FiCloud ’21)
Federica Filippini, Danilo Ardagna, Marco Lattuada, Edoardo Amaldi, Maciek Riedl, Katarzyna Materka, Paweł Skrzypek, Michele Ciavotta, Fabrizio Magugliani, and Marco Cicala. 2021 · 2021
Later among the works it cites.
Chronus: A Novel Deadline-aware Scheduler for Deep Learning Training Jobs. In Proceedings of the ACM Symposium on Cloud Computing (SoCC ’21)
Wei Gao, Zhisheng Ye, Peng Sun, Yonggang Wen, and Tianwei Zhang. 2021 · 2021
Later among the works it cites.
A Survey of Quantization Methods for Efficient Neural Network Inference
Amir Gholami, Sehoon Kim, Zhen Dong, Zhewei Yao, Michael W. Mahoney, and Kurt Keutzer. 2021 · 2021
Later among the works it cites.
Liquid: Intelligent Resource Estimation and Network-Efficient Scheduling for Deep Learning Jobs on Distributed GPU Clusters
Rong Gu, Yuquan Chen, Shuai Liu, Haipeng Dai, Guihai Chen, Kai Zhang, Yang Che, and Yihua Huang. 2021 · 2021
Later among the works it cites.
Cocktail: Leveraging Ensemble Learning for Optimized Model Serving in Public Cloud
Jashwant Raj Gunasekaran, Cyan Subhra Mishra, Prashanth Thinakaran, Mahmut Taylan Kandemir, and Chita R. Das. 2021 · 2021
Later among the works it cites.
Characterization and Prediction of Deep Learning Workloads in Large-Scale GPU Datacenters. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC ’21)
Qinghao Hu, Peng Sun, Shengen Yan, Yonggang Wen, and Tianwei Zhang. 2021 · 2021
Later among the works it cites.
Elastic Resource Sharing for Distributed Deep Learning. In 18th USENIX Symposium on Networked Systems Design and Implementation (NSDI ’21)
Changho Hwang, Taehyun Kim, Sunghyun Kim, Jinwoo Shin, and KyoungSoo Park. 2021 · 2021
Later among the works it cites.
AMPS-Inf: Automatic Model Partitioning for Serverless Inference with Cost Efficiency. In 50th International Conference on Parallel Processing (ICPP 2021)
Jananie Jarachanthan, Li Chen, Fei Xu, and Bo Li. 2021 · 2021
Later among the works it cites.
Doing more by doing less: how structured partial backpropagation improves deep learning clusters. In Proceedings of the 2nd ACM International Workshop on Distributed Machine Learning
Adarsh Kumar, Kausik Subramanian, Shivaram Venkataraman, and Aditya Akella. 2021 · 2021
Later among the works it cites.
Pruning and quantization for deep neural network acceleration: A survey
Tailin Liang, John Glossner, Lei Wang, Shaobo Shi, and Xiaotong Zhang. 2021 · 2021
Later among the works it cites.
Privacy Budget Scheduling. In 15th USENIX Symposium on Operating Systems Design and Implementation (OSDI ’21)
Tao Luo, Mingen Pan, Pierre Tholoniat, Asaf Cidon, Roxana Geambasu, and Mathias Lécuyer. 2021 · 2021
Later among the works it cites.
Scalable Deep Learning on Distributed Infrastructures: Challenges, Techniques, and Tools
Ruben Mayer and Hans-Arno Jacobsen. 2021 · 2021
Later among the works it cites.
Energy-aware Task Scheduling with Deadline Constraint in DVFS-enabled Heterogeneous Clusters
Xinxin Mei, Qiang Wang, Xiaowen Chu, Hai Liu, Yiu-Wing Leung, and Zongpeng Li. 2021 · 2021
Later among the works it cites.
Interference-Aware Scheduling for Inference Serving. In Proceedings of the 1st Workshop on Machine Learning and Systems (EuroMLSys ’21)
Daniel Mendoza, Francisco Romero, Qian Li, Neeraja J. Yadwadkar, and Christos Kozyrakis. 2021 · 2021
Later among the works it cites.
RubberBand: Cloud-Based Hyperparameter Tuning. In Proceedings of the Sixteenth European Conference on Computer Systems (EuroSys ’21)
Ujval Misra, Richard Liaw, Lisa Dunlap, Romil Bhardwaj, Kirthevasan Kandasamy, Joseph E. Gonzalez, Ion Stoica, and Alexey Tumanov. 2021 · 2021
Later among the works it cites.
Solving Large-Scale Granular Resource Allocation Problems Efficiently with POP. In Proceedings of the ACM SIGOPS 28th Symposium on Operating Systems Principles (SOSP ’21)
Deepak Narayanan, Fiodar Kazhamiaka, Firas Abuzaid, Peter Kraft, Akshay Agrawal, Srikanth Kandula, Stephen Boyd, and Matei Zaharia. 2021 · 2021
Later among the works it cites.
PieSlicer: Dynamically Improving Response Time for Cloud-based CNN Inference. In Proceedings of the ACM/SPEC International Conference on Performance Engineering (ICPE ’21)
Samuel S. Ogden, Xiangnan Kong, and Tian Guo. 2021 · 2021
Later among the works it cites.
Communication optimization strategies for distributed deep neural network training: A survey
Shuo Ouyang, Dezun Dong, Yemao Xu, and Liquan Xiao. 2021 · 2021
Later among the works it cites.
DL2: A Deep Learning-Driven Scheduler for Deep Learning Clusters
Yanghua Peng, Yixin Bao, Yangrui Chen, Chuan Wu, Chen Meng, and Wei Lin. 2021 · 2021
Later among the works it cites.
Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep Learning. In 15th USENIX Symposium on Operating Systems Design and Implementation (OSDI ’21)
Aurick Qiao, Sang Keun Choe, Suhas Jayaram Subramanya, Willie Neiswanger, Qirong Ho, Hao Zhang, Gregory R. Ganger, and Eric P. Xing. 2021 · 2021
Later among the works it cites.
INFaaS: Automated Model-less Inference Serving. In 2021 USENIX Annual Technical Conference (USENIX ATC ’21)
Francisco Romero, Qian Li, Neeraja J. Yadwadkar, and Christos Kozyrakis. 2021 · 2021
Later among the works it cites.
SLO-Aware Inference Scheduler for Heterogeneous Processors in Edge Platforms
Wonik Seo, Sanghoon Cha, Yeonjae Kim, Jaehyuk Huh, and Jongse Park. 2021 · 2021
Later among the works it cites.
A GPU Scheduling Framework to Accelerate Hyper-Parameter Optimization in Deep Learning Clusters
Jaewon Son, Yonghyuk Yoo, Khu-rai Kim, Youngjae Kim, Kwonyong Lee, and Sungyong Park. 2021 · 2021
Later among the works it cites.
Serving DNN Models with Multi-Instance GPUs: A Case of the Reconfigurable Machine Scheduling Problem
Cheng Tan, Zhichao Li, Jian Zhang, Yu Cao, Sikai Qi, Zherui Liu, Yibo Zhu, and Chuanxiong Guo. 2021 · 2021
Later among the works it cites.
Dynamic GPU Energy Optimization for Machine Learning Training Workloads
Farui Wang, Weizhe Zhang, Shichao Lai, Meng Hao, and Zheng Wang. 2021c · 2021
Later among the works it cites.
ASTRAEA: A Fair Deep Learning Scheduler for Multi-tenant GPU Clusters
Zhisheng Ye, Peng Sun, Wei Gao, Tianwei Zhang, Xiaolin Wang, Shengen Yan, and Yingwei Luo. 2021 · 2021
Later among the works it cites.
A Survey of Large-Scale Deep Learning Serving System Optimization: Challenges and Opportunities
Fuxun Yu, Di Wang, Longfei Shangguan, Minjia Zhang, Xulong Tang, Chenchen Liu, and Xiang Chen. 2021c · 2021
Later among the works it cites.
Gillis: Serving Large Neural Networks in Serverless Functions with Automatic Model Partitioning. In 2021 IEEE 41st International Conference on Distributed Computing Systems (ICDCS)
Minchen Yu, Zhifeng Jiang, Hok Chun Ng, Wei Wang, Ruichuan Chen, and Bo Li. 2021a · 2021
Later among the works it cites.
Multi Model Server: a tool for serving neural net models for inference
2022 · 2022
Closest in time.
NVIDIA A100
2022 · 2022
Closest in time.
NVIDIA Multi-Instance GPU
2022 · 2022
Closest in time.
NVIDIA Multi-Process Service
2022 · 2022
Closest in time.
Aryl: An Elastic Cluster Scheduler for Deep Learning
Jiamin Li, Hong Xu, Yibo Zhu, Zherui Liu, Chuanxiong Guo, and Cong Wang. 2022 · 2022
Closest in time.
Synergy: Looking Beyond GPUs for DNN Scheduling on Multi-Tenant Clusters. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI ’22)
Jayashree Mohan, Amar Phanishayee, Janardhan Kulkarni, and Vijay Chidambaram. 2022 · 2022
Closest in time.
Singularity: Planet-Scale, Preemptible, Elastic Scheduling of AI Workloads
Dharma Shukla, Muthian Sivathanu, Srinidhi Viswanatha, Bhargav Gulavani, Rimma Nehme, Amey Agrawal, Chen Chen, Nipun Kwatra, Ramachandran Ramjee, Pankaj Sharma, et al · 2022
Closest in time.
MLaaS in the Wild: Workload Analysis and Scheduling in Large-Scale Heterogeneous GPU Clusters. In 19th USENIX Symposium on Networked Systems Design and Implementation (NSDI ’22)
Qizhen Weng, Wencong Xiao, Yinghao Yu, Wei Wang, Cheng Wang, Jian He, Yong Li, Liping Zhang, Wei Lin, and Yu Ding. 2022 · 2022
Closest in time.
Elastic Deep Learning in Multi-Tenant GPU Clusters
Yidi Wu, Kaihao Ma, Xiao Yan, Zhi Liu, Zhenkun Cai, Yuzhen Huang, James Cheng, Han Yuan, and Fan Yu. 2022 · 2022
Closest in time.
Horus: Interference-Aware and Prediction-Based Scheduling in Deep Learning Systems
Gingfung Yeung, Damian Borowiec, Renyu Yang, Adrian Friday, Richard Harper, and Peter Garraghan. 2022 · 2022
Closest in time.
A Survey of Multi-Tenant Deep Learning Inference on GPU
Fuxun Yu, Di Wang, Longfei Shangguan, Minjia Zhang, Chenchen Liu, and Xiang Chen. 2022b · 2022
Closest in time.
GADGET: Online Resource Optimization for Scheduling Ring-All-Reduce Learning Jobs
Menglu Yu, Ye Tian, Bo Ji, Chuan Wu, Hridesh Rajan, and Jia Liu. 2022a · 2022
Closest in time.
Online Scheduling Algorithm for Heterogeneous Distributed Machine Learning Jobs
Ruiting Zhou, Jinlong Pang, Qin Zhang, Chuan Wu, Lei Jiao, Yi Zhong, and Zongpeng Li. 2022 · 2022
Closest in time.