Fetching the paper…
Reading the bibliography…
Serverless computing offers a compelling cloud model for online inference services.
A fast storage allocator
Kenneth C. Knowlton · 1965
Earlier work this paper cites.
rcuda: Reducing the number of gpu-based accelerators in high performance clusters
José Duato, Antonio J. Peña, Federico Silla, Rafael Mayo, and Enrique S. Quintana-Ortí · 2010
Earlier work this paper cites.
A gpgpu transparent virtualization component for high performance computing clouds
Giulio Giunta, Raffaele Montella, Giuseppe Agrillo, and Giuseppe Coviello · 2010
Earlier work this paper cites.
Cu2rcu: Towards the complete rcuda remote gpu virtualization and sharing solution
C. Reaño, A. J. Peña, F. Silla, J. Duato, R. Mayo, and E. S. Quintana-Ortí · 2012
Earlier work this paper cites.
vDNN: Virtualized deep neural networks for scalable, memory-efficient neural network design
Minsoo Rhu, Natalia Gimelshein, Jason Clemons, Arslan Zulfiqar, and Stephen W. Keckler · 2016
Earlier work this paper cites.
Topology-aware gpu scheduling for learning workloads in cloud environments
Marcelo Amaral, Jordà Polo, David Carrera, Seetharami Seelam, and Malgorzata Steinder · 2017
Earlier work this paper cites.
Clipper: A low-latency online prediction serving system
Daniel Crankshaw, Xin Wang, Giulio Zhou, Michael J Franklin, Joseph E Gonzalez, and Ion Stoica · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
PRETZEL: Opening the black box of machine learning prediction serving systems
Yunseong Lee, Alberto Scolari, Byung-Gon Chun, Marco Domenico Santambrogio, Markus Weimer, and Matteo Interlandi · 2018
Earlier work this paper cites.
https://developer.download.nvidia.com/video/gputechconf/gtc/2019/presentation/s9727-memory-management-on-modern-gpu-architectures.pdf
Memory Management on Modern GPU Architectures · 2019
Earlier work this paper cites.
Parity models: erasure-coded resilience for prediction serving systems
Jack Kosaian, K. V. Rashmi, and Shivaram Venkataraman · 2019
Earlier work this paper cites.
Nexus: a GPU cluster engine for accelerating DNN-based video analysis
Haichen Shen, Lequn Chen, Yuchen Jin, Liangyu Zhao, Bingyu Kong, Matthai Philipose, Arvind Krishnamurthy, and Ravi Sundaram · 2019
Earlier work this paper cites.
MArk: Exploiting cloud services for cost-effective, SLO-aware machine learning inference serving
Chengliang Zhang, Minchen Yu, Wei Wang, and Feng Yan · 2019
Earlier work this paper cites.
Firecracker: Lightweight virtualization for serverless applications
Alexandru Agache, Marc Brooker, Alexandra Iordache, Anthony Liguori, Rolf Neugebauer, Phil Piwonka, and Diana-Maria Popa · 2020
Earlier work this paper cites.
BATCH: Machine learning inference serving on serverless platforms with adaptive batching
Ahsan Ali, Riccardo Pinciroli, Feng Yan, and Evgenia Smirni · 2020
Earlier work this paper cites.
PipeSwitch: Fast pipelined context switching for deep learning applications
Zhihao Bai, Zhen Zhang, Yibo Zhu, and Xin Jin · 2020
Earlier work this paper cites.
Serving DNNs like clockwork: Performance predictability from the bottom up
Arpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao, Antoine Kaufmann, Ymir Vigfusson, and Jonathan Mace · 2020
Earlier work this paper cites.
SwapAdvisor: Pushing deep learning beyond the GPU memory limit via smart swapping
Chien-Chin Huang, Gu Jin, and Jinyang Li · 2020
Cited alongside, same era.
A unified architecture for accelerating distributed dnn training in heterogeneous gpu/cpu clusters
Yimin Jiang, Yibo Zhu, Chang Lan, Bairen Yi, Yong Cui, and Chuanxiong Guo · 2020
Cited alongside, same era.
Batch-aware unified memory management in GPUs for irregular workloads
Hyojong Kim, Jaewoong Sim, Prasun Gera, Ramyad Hadidi, and Hyesoon Kim · 2020
Cited alongside, same era.
Capuchin: Tensor-based GPU memory management for deep learning
Xuan Peng, Xuanhua Shi, Hulin Dai, Hai Jin, Weiliang Ma, Qian Xiong, Fan Yang, and Xuehai Qian · 2020
Cited alongside, same era.
Serverless in the wild: Characterizing and optimizing the serverless workload at a large cloud provider
Mohammad Shahrad, Rodrigo Fonseca, Iñigo Goiri, Gohar Chaudhry, Paul Batum, Jason Cooke, Eduardo Laureano, Colby Tresness, Mark Russinovich, and Ricardo Bianchini · 2020
Cited alongside, same era.
Dgsf: Disaggregated gpus for serverless functions
Henrique Fingler, Zhiting Zhu, Esther Yoon, Zhipeng Jia, Emmett Witchel, and Christopher J. Rossbach · 2022
Later among the works it cites.
Microsecond-scale preemption for concurrent GPU-accelerated DNN inferences
Mingcong Han, Hanze Zhang, Rong Chen, and Haibo Chen · 2022
Later among the works it cites.
Owl: Performance-aware scheduling for resource-efficient function-as-a-service cloud
Huangshi Tian, Suyi Li, Ao Wang, Wei Wang, Tianlong Wu, and Haoran Yang · 2022
Later among the works it cites.
INFless: a native serverless system for low-latency, high-throughput inference
Yanan Yang, Laiping Zhao, Yiming Li, Huanyu Zhang, Jie Li, Mingyang Zhao, Xingzhen Chen, and Keqiu Li · 2022
Later among the works it cites.
Fast and efficient model serving using multi-gpus with direct-host-access
Jinwoo Jeong, Seungsu Baek, and Jeongseob Ahn · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
AntMan: Dynamic scaling on GPU clusters for deep learning
Wencong Xiao, Shiru Ren, Yong Li, Yang Zhang, Pengyang Hou, Zhi Li, Yihui Feng, Wei Lin, and Yangqing Jia · 2020
Cited alongside, same era.
AvA: Accelerated virtualization of accelerators
Hangchen Yu, Arthur Michener Peters, Amogh Akshintala, and Christopher J. Rossbach · 2020
Cited alongside, same era.
Salus: Fine-grained GPU sharing primitives for deep learning applications
Peifeng Yu and Mosharaf Chowdhury · 2020
Cited alongside, same era.
Nightcore: Efficient and scalable serverless computing for latency-sensitive, interactive microservices
Zhipeng Jia and Emmett Witchel · 2021
Cited alongside, same era.
INFaaS: Automated model-less inference serving
Francisco Romero, Qian Li, Neeraja J Yadwadkar, and Christos Kozyrakis · 2021
Cited alongside, same era.
Benchmarking, analysis, and optimization of serverless function snapshots
Dmitrii Ustiugov, Plamen Petrov, Marios Kogias, Edouard Bugnion, and Boris Grot · 2021
Cited alongside, same era.
FaaSNet: Scalable and fast provisioning of custom serverless container runtimes at alibaba cloud function compute
Ao Wang, Shuai Chang, Huangshi Tian, Hongqi Wang, Haoran Yang, Huiba Li, Rui Du, and Yue Cheng · 2021
Cited alongside, same era.
Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Daniel Y. Fu, Zhiqiang Xie, Beidi Chen, Clark Barrett, Joseph E. Gonzalez, Percy Liang, Christopher Ré, Ion Stoica, and Ce Zhang · 2023
Closest in time.
Following the data, not the function: Rethinking function orchestration in serverless computing
Minchen Yu, Tingjia Cao, Wei Wang, and Ruichuan Chen · 2023
Closest in time.
SHEPHERD: Serving DNNs in the wild
Hong Zhang, Yupeng Tang, Anurag Khandelwal, and Ion Stoica · 2023
Closest in time.
ServerlessLLM: Locality-enhanced serverless inference for large language models
Yao Fu, Leyang Xue, Yeqi Huang, Andrei-Octavian Brabete, Dmitrii Ustiugov, Yuvraj Patel, and Luo Mai · 2024
Closest in time.
PARALLELGPUOS: A concurrent OS-level GPU checkpoint and restore system using validated speculation
Zhuobin Huang, Xingda Wei, Yingyi Hao, Rong Chen, Mingcong Han, Jinyu Gu, and Haibo Chen · 2024
Closest in time.
Orion: Interference-aware, fine-grained gpu sharing for ml applications
Foteini Strati, Xianzhe Ma, and Ana Klimovic · 2024
Closest in time.
Faastube: Optimizing gpu-oriented data transfer for serverless computing
Hao Wu, Junxiao Deng, Minchen Yu, Yue Yu, Yaochen Liu, Hao Fan, Song Wu, and Wei Wang · 2024
Closest in time.
StreamBox: A lightweight GPU SandBox for serverless inference workflow
Hao Wu, Yue Yu, Junxiao Deng, Shadi Ibrahim, Song Wu, Hao Fan, Ziyue Cheng, and Hai Jin · 2024
Closest in time.
DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving
Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang · 2024
Closest in time.
Dilu: Enabling GPU resourcing-on-demand for serverless DL serving via introspective elasticity
Cunchi Lv, Xiao Shi, Zhengyu Lei, Jinyue Huang, Wenting Tan, Xiaohui Zheng, and Xiaofang Zhao · 2025
Closest in time.
Medusa: Accelerating serverless LLM inference with materialization
Shaoxun Zeng, Minhui Xie, Shiwei Gao, Youmin Chen, and Youyou Lu · 2025
Closest in time.