Fetching the paper…
Reading the bibliography…
Serverless computing has gained significant traction for machine learning inference applications, which are often deployed as serverless workflows consisting of multiple CPU and GPU functions with data dependency.
TensorFlow: a system for large-scale machine learning. In Proceedings of the 12th USENIX Conference on Operating Systems Design and Implementation (Savannah, GA, USA) (OSDI’16) . USENIX Association, USA, 265–283
Martín Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, Manjunath Kudlur, Josh Levenberg, Rajat Monga, Sherry Moore, Derek G. Murray, Benoit Steiner, Paul Tucker, Vijay Vasudevan, Pete Warden, Martin Wicke, Yuan Yu, and Xiaoqiang Zheng. 2016 · 2016
Earlier work this paper cites.
PyTorch: An Imperative Style, High-Performance Deep Learning Library. In Advances in Neural Information Processing Systems , H. Wallach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett (Eds.), Vol. 32. Curran Associates, Inc
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, Alban Desmaison, Andreas Kopf, Edward Yang, Zachary DeVito, Martin Raison, Alykhan Tejani, Sasank Chilamkurthy, Benoit Steiner, Lu Fang, Junjie Bai, and Soumith Chintala. 2019 · 2019
Earlier work this paper cites.
Nexus: a GPU cluster engine for accelerating DNN-based video analysis. In Proceedings of the 27th ACM Symposium on Operating Systems Principles (Huntsville, Ontario, Canada) (SOSP ’19) . Association for Computing Machinery, New York, NY, USA, 322–337
Haichen Shen, Lequn Chen, Yuchen Jin, Liangyu Zhao, Bingyu Kong, Matthai Philipose, Arvind Krishnamurthy, and Ravi Sundaram. 2019 · 2019
Earlier work this paper cites.
BATCH: Machine Learning Inference Serving on Serverless Platforms with Adaptive Batching. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis . 1–15
Ahsan Ali, Riccardo Pinciroli, Feng Yan, and Evgenia Smirni. 2020 · 2020
Earlier work this paper cites.
Balancing efficiency and fairness in heterogeneous GPU clusters for deep learning. In Proceedings of the Fifteenth European Conference on Computer Systems (Heraklion, Greece) (EuroSys ’20) . Association for Computing Machinery, New York, NY, USA, Article 1, 16 pages
Shubham Chaudhary, Ramachandran Ramjee, Muthian Sivathanu, Nipun Kwatra, and Srinidhi Viswanatha. 2020 · 2020
Earlier work this paper cites.
InferLine: latency-aware provisioning and scaling for prediction serving pipelines. In Proceedings of the 11th ACM Symposium on Cloud Computing (Virtual Event, USA) (SoCC ’20) . Association for Computing Machinery, New York, NY, USA, 477–491
Daniel Crankshaw, Gur-Eyal Sela, Xiangxi Mo, Corey Zumar, Ion Stoica, Joseph Gonzalez, and Alexey Tumanov. 2020 · 2020
Earlier work this paper cites.
Serverless in the Wild: Characterizing and Optimizing the Serverless Workload at a Large Cloud Provider. In 2020 USENIX Annual Technical Conference (USENIX ATC 20) . USENIX Association, 205–218
Mohammad Shahrad, Rodrigo Fonseca, Inigo Goiri, Gohar Chaudhry, Paul Batum, Jason Cooke, Eduardo Laureano, Colby Tresness, Mark Russinovich, and Ricardo Bianchini. 2020 · 2020
Earlier work this paper cites.
Blink: Fast and Generic Collectives for Distributed ML. In Proceedings of Machine Learning and Systems , I. Dhillon, D. Papailiopoulos, and V. Sze (Eds.), Vol. 2. 172–186
Guanhua Wang, Shivaram Venkataraman, Amar Phanishayee, Nikhil Devanur, Jorgen Thelin, and Ion Stoica. 2020 · 2020
Earlier work this paper cites.
Synthesizing optimal collective algorithms. In Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (Virtual Event, Republic of Korea) (PPoPP ’21) . Association for Computing Machinery, New York, NY, USA, 62–75
Zixian Cai, Zhengyang Liu, Saeed Maleki, Madanlal Musuvathi, Todd Mytkowicz, Jacob Nelson, and Olli Saarikivi. 2021 · 2021
Earlier work this paper cites.
Scrooge: A Cost-Effective Deep Learning Inference System. In Proceedings of the ACM Symposium on Cloud Computing (Seattle, WA, USA) (SoCC ’21) . Association for Computing Machinery, New York, NY, USA, 624–638
Yitao Hu, Rajrup Ghosh, and Ramesh Govindan. 2021 · 2021
Earlier work this paper cites.
Nightcore: efficient and scalable serverless computing for latency-sensitive, interactive microservices. In Proceedings of the 26th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (Virtual, USA) (ASPLOS ’21) . Association for Computing Machinery, New York, NY, USA, 152–166
Zhipeng Jia and Emmett Witchel. 2021 · 2021
Earlier work this paper cites.
SONIC: Application-aware Data Passing for Chained Serverless Applications. In 2021 USENIX Annual Technical Conference (USENIX ATC 21) . USENIX Association, 285–301
Ashraf Mahgoub, Karthick Shankar, Subrata Mitra, Ana Klimovic, Somali Chaterji, and Saurabh Bagchi. 2021 · 2021
Earlier work this paper cites.
MAPA: multi-accelerator pattern allocation policy for multi-tenant GPU servers. In Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (St. Louis, Missouri) (SC ’21) . Association for Computing Machinery, New York, NY, USA, Article 99, 14 pages
Kiran Ranganath, Joshua D. Suetterlein, Joseph B. Manzano, Shuaiwen Leon Song, and Daniel Wong. 2021 · 2021
Earlier work this paper cites.
Llama: A Heterogeneous & Serverless Framework for Auto-Tuning Video Analytics Pipelines. In Proceedings of the ACM Symposium on Cloud Computing (Seattle, WA, USA) (SoCC ’21) . Association for Computing Machinery, New York, NY, USA, 1–17
Francisco Romero, Mark Zhao, Neeraja J. Yadwadkar, and Christos Kozyrakis. 2021 · 2021
Earlier work this paper cites.
Memory Harvesting in Multi-GPU Systems with Hierarchical Unified Virtual Memory. In 2022 USENIX Annual Technical Conference (USENIX ATC 22) . USENIX Association, Carlsbad, CA, 625–638
Sangjin Choi, Taeksoo Kim, Jinwoo Jeong, Rachata Ausavarungnirun, Myeongjae Jeon, Youngjin Kwon, and Jeongseob Ahn. 2022a · 2022
Cited alongside, same era.
Serving Heterogeneous Machine Learning Models on Multi-GPU Servers with Spatio-Temporal Sharing. In 2022 USENIX Annual Technical Conference (USENIX ATC 22) . USENIX Association, Carlsbad, CA, 199–216
Seungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park, Youngjin Kwon, and Jaehyuk Huh. 2022b · 2022
Cited alongside, same era.
DGSF: Disaggregated GPUs for Serverless Functions. In 2022 IEEE International Parallel and Distributed Processing Symposium (IPDPS) . 739–750
Henrique Fingler, Zhiting Zhu, Esther Yoon, Zhipeng Jia, Emmett Witchel, and Christopher J. Rossbach. 2022 · 2022
Cited alongside, same era.
Cocktail: A Multidimensional Optimization for Model Serving in Cloud. In 19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22) . USENIX Association, Renton, WA, 1041–1057
Jashwant Raj Gunasekaran, Cyan Subhra Mishra, Prashanth Thinakaran, Bikash Sharma, Mahmut Taylan Kandemir, and Chita R. Das. 2022 · 2022
AdaInf: Data Drift Adaptive Scheduling for Accurate and SLO-guaranteed Multiple-Model Inference Serving at Edge Servers. In Proceedings of the ACM SIGCOMM 2023 Conference (<conf-loc>, <city>New York</city>, <state>NY</state>, <country>USA</country>, </conf-loc>) (ACM SIGCOMM ’23) . Association for Computing Machinery, New York, NY, USA, 473–485
Sudipta Saha Shubha and Haiying Shen. 2023 · 2023
Later among the works it cites.
Transparent GPU Sharing in Container Clouds for Deep Learning Workloads. In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23) . USENIX Association, Boston, MA, 69–85
Bingyang Wu, Zili Zhang, Zhihao Bai, Xuanzhe Liu, and Xin Jin. 2023 · 2023
Later among the works it cites.
FaaSwap: SLO-Aware, GPU-Efficient Serverless Inference via Model Swapping
Minchen Yu, Ao Wang, Dong Chen, Haoxuan Yu, Xiaonan Luo, Zhuohao Li, Wei Wang, Ruichuan Chen, Dapeng Nie, and Haoran Yang. 2023b · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Microsecond-scale Preemption for Concurrent GPU-accelerated DNN Inferences. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22) . USENIX Association, Carlsbad, CA, 539–558
Mingcong Han, Hanze Zhang, Rong Chen, and Haibo Chen. 2022 · 2022
Cited alongside, same era.
Tetris: Memory-efficient Serverless Inference through Tensor Sharing. In 2022 USENIX Annual Technical Conference (USENIX ATC 22) . USENIX Association, Carlsbad, CA
Jie Li, Laiping Zhao, Yanan Yang, Kunlin Zhan, and Keqiu Li. 2022b · 2022
Cited alongside, same era.
MLaaS in the Wild: Workload Analysis and Scheduling in Large-Scale Heterogeneous GPU Clusters. In 19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22) . USENIX Association, Renton, WA, 945–960
Qizhen Weng, Wencong Xiao, Yinghao Yu, Wei Wang, Cheng Wang, Jian He, Yong Li, Liping Zhang, Wei Lin, and Yu Ding. 2022 · 2022
Cited alongside, same era.
INFless: a native serverless system for low-latency, high-throughput inference. In Proceedings of the 27th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (Lausanne, Switzerland) (ASPLOS ’22) . Association for Computing Machinery, New York, NY, USA, 768–781
Yanan Yang, Laiping Zhao, Yiming Li, Huanyu Zhang, Jie Li, Mingyang Zhao, Xingzhen Chen, and Keqiu Li. 2022 · 2022
Cited alongside, same era.
Astraea: towards QoS-aware and resource-efficient multi-stage GPU services. In Proceedings of the 27th ACM International Conference on Architectural Support for Programming Languages and Operating Systems (Lausanne, Switzerland) (ASPLOS ’22) . Association for Computing Machinery, New York, NY, USA, 570–582
Wei Zhang, Quan Chen, Kaihua Fu, Ningxin Zheng, Zhiyi Huang, Jingwen Leng, and Minyi Guo. 2022 · 2022
Cited alongside, same era.
Boggart: Towards General-Purpose Acceleration of Retrospective Video Analytics. In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23) . USENIX Association, Boston, MA, 933–951
Neil Agarwal and Ravi Netravali. 2023 · 2023
Cited alongside, same era.
Fast and Efficient Model Serving Using Multi-GPUs with Direct-Host-Access. In Proceedings of the Eighteenth European Conference on Computer Systems (Rome, Italy) (EuroSys ’23) . Association for Computing Machinery, New York, NY, USA, 249–265
Jinwoo Jeong, Seungsu Baek, and Jeongseob Ahn. 2023 · 2023
Cited alongside, same era.
DeepUM: Tensor Migration and Prefetching in Unified Memory. In Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 (Vancouver, BC, Canada) (ASPLOS 2023) . Association for Computing Machinery, New York, NY, USA, 207–221
Jaehoon Jung, Jinpyo Kim, and Jaejin Lee. 2023 · 2023
Cited alongside, same era.
Hong Zhang, Yupeng Tang, Anurag Khandelwal, and Ion Stoica. 2023 · 2023
Later among the works it cites.
AQUATOPE: QoS-and-Uncertainty-Aware Resource Management for Multi-stage Serverless Workflows. In Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1 (Vancouver, BC, Canada) (ASPLOS 2023) . Association for Computing Machinery, New York, NY, USA, 1–14
Zhuangzhuang Zhou, Yanqi Zhang, and Christina Delimitrou. 2022 · 2023
Later among the works it cites.
ServerlessLLM: Low-Latency Serverless Inference for Large Language Models. In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) . USENIX Association, Santa Clara, CA, 135–153
Yao Fu, Leyang Xue, Yeqi Huang, Andrei-Octavian Brabete, Dmitrii Ustiugov, Yuvraj Patel, and Luo Mai. 2024 · 2024
Closest in time.
GMLake: Efficient and Transparent GPU Memory Defragmentation for Large-scale DNN Training with Virtual Memory Stitching. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 (<conf-loc>, <city>La Jolla</city>, <state>CA</state>, <country>USA</country>, </conf-loc>) (ASPLOS ’24) . Association for Computing Machinery, New York, NY, USA, 450–466
Cong Guo, Rui Zhang, Jiale Xu, Jingwen Leng, Zihan Liu, Ziyu Huang, Minyi Guo, Hao Wu, Shouren Zhao, Junping Zhao, and Ke Zhang. 2024 · 2024
Closest in time.
TCCL: Discovering Better Communication Paths for PCIe GPU Clusters. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3 (La Jolla, CA, USA) (ASPLOS ’24) . Association for Computing Machinery, New York, NY, USA, 999–1015
Heehoon Kim, Junyeol Ryu, and Jaejin Lee. 2024 · 2024
Closest in time.
DataFlower: Exploiting the Data-flow Paradigm for Serverless Workflow Orchestration. In Proceedings of the 28th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 4 (<conf-loc>, <city>Vancouver</city>, <state>BC</state>, <country>Canada</country>, </conf-loc>) (ASPLOS ’23) . Association for Computing Machinery, New York, NY, USA, 57–72
Zijun Li, Chuhao Xu, Quan Chen, Jieru Zhao, Chen Chen, and Minyi Guo. 2024 · 2024
Closest in time.
FUYAO: DPU-enabled Direct Data Transfer for Serverless Computing. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3 (<conf-loc>, <city>La Jolla</city>, <state>CA</state>, <country>USA</country>, </conf-loc>) (ASPLOS ’24) . Association for Computing Machinery, New York, NY, USA, 431–447
Guowei Liu, Laiping Zhao, Yiming Li, Zhaolin Duan, Sheng Chen, Yitao Hu, Zhiyuan Su, and Wenyu Qu. 2024 · 2024
Closest in time.
Serialization/Deserialization-free State Transfer in Serverless Workflows. In Proceedings of the Nineteenth European Conference on Computer Systems (Athens, Greece) (EuroSys ’24) . Association for Computing Machinery, New York, NY, USA, 132–147
Fangming Lu, Xingda Wei, Zhuobin Huang, Rong Chen, Minyu Wu, and Haibo Chen. 2024 · 2024
Closest in time.
Enhancing Intra-Node GPU-to-GPU Performance in MPI+UCX through Multi-Path Communication. In Proceedings of the 3rd International Workshop on Extreme Heterogeneity Solutions (Edinburgh, United Kingdom) (ExHET ’24) . Association for Computing Machinery, New York, NY, USA, 9–14
Amirhossein Sojoodi, Yiltan H. Temucin, and Ahmad Afsahi. 2024 · 2024
Closest in time.
StreamBox: A Lightweight GPU SandBox for Serverless Inference Workflow. In 2024 USENIX Annual Technical Conference (USENIX ATC 24) . USENIX Association, Santa Clara, CA, 59–73
Hao Wu, Yue Yu, Junxiao Deng, Shadi Ibrahim, Song Wu, Hao Fan, Ziyue Cheng, and Hai Jin. 2024 · 2024
Closest in time.