Fetching the paper…
Reading the bibliography…
Large language model (LLM) serving is becoming an increasingly important workload for cloud providers.
Some methods for classification and analysis of multivariate observations
J MacQueen · 1967
Earlier work this paper cites.
Congestion avoidance and control
Van Jacobson · 1988
Earlier work this paper cites.
Database architecture optimized for the new bottleneck: Memory access
Peter A Boncz, Stefan Manegold, Martin L Kersten, et al · 1999
Earlier work this paper cites.
Understanding the Linux Kernel: from I/O ports to process management
Daniel P Bovet and Marco Cesati · 2005
Earlier work this paper cites.
Morpheus: Towards automated SLO for enterprise clusters
Sangeetha Abdu Jyothi, Carlo Curino, Ishai Menache, Shravan Matthur Narayanamurthy, Alexey Tumanov, Jonathan Yaniv, Ruslan Mavlyutov, Inigo Goiri, Subru Krishnan, Janardhan Kulkarni, et al · 2016
Earlier work this paper cites.
Clipper: A low-latency online prediction serving system
Daniel Crankshaw, Xin Wang, Guilio Zhou, Michael J Franklin, Joseph E Gonzalez, and Ion Stoica · 2017
Earlier work this paper cites.
TensorFlow-Serving: Flexible, high-performance ML serving
Christopher Olston, Fangwei Li, Jeremiah Harmsen, Jordan Soyke, Kiril Gorovoy, Li Lao, Noah Fiedel, Sukriti Ramesh, and Vinu Rajashekhar · 2017
Earlier work this paper cites.
MArk: Exploiting cloud services for cost-effective, SLO-aware machine learning inference serving
Chengliang Zhang, Minchen Yu, Wei Wang, and Feng Yan · 2019
Earlier work this paper cites.
InferLine: Latency-aware provisioning and scaling for prediction serving pipelines
Daniel Crankshaw, Gur-Eyal Sela, Xiangxi Mo, Corey Zumar, Ion Stoica, Joseph Gonzalez, and Alexey Tumanov · 2020
Earlier work this paper cites.
Serving DNNs like clockwork: Performance predictability from the bottom up
Arpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao, Antoine Kaufmann, Ymir Vigfusson, and Jonathan Mace · 2020
Earlier work this paper cites.
FIRM: An intelligent fine-grained resource management framework for SLO-oriented microservices
Haoran Qiu, Subho S Banerjee, Saurabh Jha, Zbigniew T Kalbarczyk, and Ravishankar K Iyer · 2020
Earlier work this paper cites.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Earlier work this paper cites.
TurboTransformers: An efficient GPU serving system for transformer models
Jiarui Fang, Yang Yu, Chengduo Zhao, and Jie Zhou · 2021
Earlier work this paper cites.
INFaaS: Automated model-less inference serving
Francisco Romero, Qian Li, Neeraja J Yadwadkar, and Christos Kozyrakis · 2021
Earlier work this paper cites.
Cocktail: A multidimensional optimization for model serving in cloud
Jashwant Raj Gunasekaran, Cyan Subhra Mishra, Prashanth Thinakaran, Bikash Sharma, Mahmut Taylan Kandemir, and Chita R Das · 2022
Cited alongside, same era.
Orca: A distributed serving system for Transformer-based generative models
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun · 2022
Cited alongside, same era.
Alpa: Automating inter-and intra-operator parallelism for distributed deep learning
Lianmin Zheng, Zhuohan Li, Hao Zhang, Yonghao Zhuang, Zhifeng Chen, Yanping Huang, Yida Wang, Yuanzhong Xu, Danyang Zhuo, Eric P Xing, et al · 2022
Cited alongside, same era.
Efficient memory management for large language model serving with PagedAttention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica · 2023
Cited alongside, same era.
Fast inference from transformers via speculative decoding
Yaniv Leviathan, Matan Kalman, and Yossi Matias · 2023
Cited alongside, same era.
ServerlessLLM: Locality-enhanced serverless inference for large language models, 2024
Yao Fu, Leyang Xue, Yeqi Huang, Andrei-Octavian Brabete, Dmitrii Ustiugov, Yuvraj Patel, and Luo Mai · 2024
Later among the works it cites.
Model tells you what to discard: Adaptive kv cache compression for LLMs, 2024
Suyu Ge, Yunan Zhang, Liyuan Liu, Minjia Zhang, Jiawei Han, and Jianfeng Gao · 2024
Later among the works it cites.
LLM inference series: 4. KV caching, a deeper look, 2024
Pierre Lienhart · 2024
Later among the works it cites.
Andes: Defining and enhancing quality-of-experience in LLM-based text streaming services
Jiachen Liu, Zhiyu Wu, Jae-Won Chung, Fan Lai, Myungjin Lee, and Mosharaf Chowdhury · 2024
Later among the works it cites.
One queue is all you need: Resolving head-of-line blocking in large language model serving
Archit Patke, Dhemath Reddy, Saurabh Jha, Haoran Qiu, Christian Pinto, Shengkun Cui, Chandra Narayanaswami, Zbigniew Kalbarczyk, and Ravishankar Iyer · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
DejaVu: Contextual sparsity for efficient LLMs at inference time
Zichang Liu, Jue Wang, Tri Dao, Tianyi Zhou, Binhang Yuan, Zhao Song, Anshumali Shrivastava, Ce Zhang, Yuandong Tian, Christopher Re, et al · 2023
Cited alongside, same era.
SpotServe: Serving generative large language models on preemptible instances
Xupeng Miao, Chunan Shi, Jiangfei Duan, Xiaoli Xi, Dahua Lin, Bin Cui, and Zhihao Jia · 2023
Cited alongside, same era.
Splitwise: Efficient generative LLM inference using phase splitting, 2023
Pratyush Patel, Esha Choukse, Chaojie Zhang, Íñigo Goiri, Aashaka Shah, Saeed Maleki, and Ricardo Bianchini · 2023
Cited alongside, same era.
FlexGen: High-throughput generative inference of large language models with a single GPU
Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher Ré, Ion Stoica, and Ce Zhang · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, et al · 2023
Cited alongside, same era.
Shepherd: Serving DNNs in the wild
Hong Zhang, Yupeng Tang, Anurag Khandelwal, and Ion Stoica · 2023
Cited alongside, same era.
APIServe: Efficient API support for large-language model inferencing
Reyna Abhyankar, Zijian He, Vikranth Srivatsa, Hao Zhang, and Yiying Zhang · 2024
Cited alongside, same era.
Later among the works it cites.
ShareGPT dataset, 2024
sharegpt · 2024
Later among the works it cites.
Llumnix: Dynamic scheduling for large language model serving
Biao Sun, Ziming Huang, Hanyu Zhao, Wencong Xiao, Xinyi Zhang, Yong Li, and Wei Lin · 2024
Later among the works it cites.
TensorRT-LLM
tensorrt · 2024
Later among the works it cites.
Text Generation Inference
tgi · 2024
Later among the works it cites.
Nvidia Triton Inference Server
triton · 2024
Later among the works it cites.
Towards efficient and reliable LLM serving: A real-world workload study, 2024
Yuxin Wang, Yuhan Chen, Zeyu Li, Zhenheng Tang, Rui Guo, Xin Wang, Qiang Wang, Amelie Chi Zhou, and Xiaowen Chu · 2024
Later among the works it cites.
The shift from models to compound AI systems
Matei Zaharia, Omar Khattab, Lingjiao Chen, Jared Quincy Davis, Heather Miller, Chris Potts, James Zou, Michael Carbin, Jonathan Frankle, Naveen Rao, and Ali Ghodsi · 2024
Later among the works it cites.
DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving, 2024
Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang · 2024
Later among the works it cites.
RelayAttention for efficient large language model serving with long system prompts, 2024
Lei Zhu, Xinjiang Wang, Wayne Zhang, and Rynson W. H. Lau · 2024
Later among the works it cites.