Fetching the paper…
Reading the bibliography…
Large language model (LLM) serving is becoming an increasingly critical workload for cloud providers.
Deep Residual Learning for Image Recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR 2016) . IEEE, Las Vegas, NV, 770–778
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016 · 2016
Earlier work this paper cites.
Morpheus: Towards automated SLO for enterprise clusters. In Proceedings of the 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 2016) . USENIX, Savannah, GA, 117–134
Sangeetha Abdu Jyothi, Carlo Curino, Ishai Menache, Shravan Matthur Narayanamurthy, Alexey Tumanov, Jonathan Yaniv, Ruslan Mavlyutov, Inigo Goiri, Subru Krishnan, Janardhan Kulkarni, et al · 2016
Earlier work this paper cites.
Clipper: A Low-Latency Online Prediction Serving System. In Proceedings of the 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 2017) . USENIX, Boston,MA, 613–627
Daniel Crankshaw, Xin Wang, Guilio Zhou, Michael J Franklin, Joseph E Gonzalez, and Ion Stoica. 2017 · 2017
Earlier work this paper cites.
TensorFlow-Serving: Flexible, High-Performance ML Serving
Christopher Olston, Fangwei Li, Jeremiah Harmsen, Jordan Soyke, Kiril Gorovoy, Li Lao, Noah Fiedel, Sukriti Ramesh, and Vinu Rajashekhar. 2017 · 2017
Earlier work this paper cites.
Recent advances in recurrent neural networks
Hojjat Salehinejad, Sharan Sankar, Joseph Barfett, Errol Colak, and Shahrokh Valaee. 2017 · 2017
Earlier work this paper cites.
Parties: QoS-aware resource partitioning for multiple interactive services. In Proceedings of the 24th International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS 2019) . ACM, Providence, RI, 107–120
Shuang Chen, Christina Delimitrou, and José F Martínez. 2019 · 2019
Earlier work this paper cites.
MArk: Exploiting Cloud Services for Cost-Effective, SLO-Aware Machine Learning Inference Serving. In Proceedings of 2019 USENIX Annual Technical Conference (ATC 2019) . USENIX, Renton, WA, 1049–1062
Chengliang Zhang, Minchen Yu, Wei Wang, and Feng Yan. 2019 · 2019
Earlier work this paper cites.
InferLine: Latency-Aware Provisioning and Scaling for Prediction Serving Pipelines. In Proceedings of the 11th ACM Symposium on Cloud Computing . Association for Computing Machinery, New York, NY, USA, 477–491
Daniel Crankshaw, Gur-Eyal Sela, Xiangxi Mo, Corey Zumar, Ion Stoica, Joseph Gonzalez, and Alexey Tumanov. 2020 · 2020
Earlier work this paper cites.
Serving DNNs like clockwork: Performance predictability from the bottom up. In Proceedings of the 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 2020) . USENIX, Berkeley, CA, 443–462
Arpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao, Antoine Kaufmann, Ymir Vigfusson, and Jonathan Mace. 2020 · 2020
Earlier work this paper cites.
Perseus: Characterizing performance and cost of multi-tenant serving for CNN models. In 2020 IEEE International Conference on Cloud Engineering (IC2E 2020) . IEEE, IEEE, Boston, MA, 66–72
Matthew LeMay, Shijian Li, and Tian Guo. 2020 · 2020
Earlier work this paper cites.
FIRM: An intelligent fine-grained resource management framework for SLO-oriented microservices. In Proceedings of The 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 2020) . USENIX, Virtual, 805–825
Haoran Qiu, Subho S Banerjee, Saurabh Jha, Zbigniew T Kalbarczyk, and Ravishankar K Iyer. 2020 · 2020
Earlier work this paper cites.
On the Opportunities and Risks of Foundation Models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al · 2021
Earlier work this paper cites.
The Big-M method with the numerical infinite M
Marco Cococcioni and Lorenzo Fiaschi. 2021 · 2021
Earlier work this paper cites.
TurboTransformers: An Efficient GPU Serving System for Transformer Models. In Proceedings of the 26th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (PPoPP 2021) . ACM, Austin, TX, 389–402
Jiarui Fang, Yang Yu, Chengduo Zhao, and Jie Zhou. 2021 · 2021
Earlier work this paper cites.
INFaaS: Automated Model-less Inference Serving. In Proceedings of 2021 USENIX Annual Technical Conference (ATC 2021) . USENIX, Virtual, 397–411
Francisco Romero, Qian Li, Neeraja J Yadwadkar, and Christos Kozyrakis. 2021 · 2021
Earlier work this paper cites.
Serving heterogeneous machine learning models on multi-GPU servers with Spatio-Temporal sharing. In 2022 USENIX Annual Technical Conference (USENIX ATC 2022) . USENIX, Carlsbad, CA, 199–216
Seungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park, Youngjin Kwon, and Jaehyuk Huh. 2022 · 2022
Earlier work this paper cites.
Opposing Effects of Response Time in Human–Chatbot Interaction: the moderating role of prior experience
Ulrich Gnewuch, Stefan Morana, Marc TP Adam, and Alexander Maedche. 2022 · 2022
Earlier work this paper cites.
Cocktail: A Multidimensional Optimization for Model Serving in Cloud. In Proceedings of the 19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 2022) . USENIX, Renton, WA, 1041–1057
Jashwant Raj Gunasekaran, Cyan Subhra Mishra, Prashanth Thinakaran, Bikash Sharma, Mahmut Taylan Kandemir, and Chita R Das. 2022 · 2022
Earlier work this paper cites.
Can foundation models wrangle your data?
Avanika Narayan, Ines Chami, Laurel Orr, Simran Arora, and Christopher Ré. 2022 · 2022
Cited alongside, same era.
Orca: A Distributed Serving System for Transformer-Based Generative Models. In Proceedings of the 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 2022) . USENIX, Carlsbad, CA, 521–538
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun. 2022 · 2022
Cited alongside, same era.
Vicuna: An Open-Source Chatbot Impressing GPT-4 with 90% ChatGPT Quality
Wei-Lin Chiang, Zhuohan Li, Zi Lin, Ying Sheng, Zhanghao Wu, Hao Zhang, Lianmin Zheng, Siyuan Zhuang, Yonghao Zhuang, Joseph E. Gonzalez, Ion Stoica, and Eric P. Xing. 2023 · 2023
Cited alongside, same era.
Mistral 7B
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al · 2023
Cited alongside, same era.
Efficient Memory Management for Large Language Model Serving with PagedAttention. In Proceedings of the 29th Symposium on Operating Systems Principles (SOSP 2023) . ACM, Koblenz, Germany, 611–626
Text Generation Inference
HuggingFace. 2024 · 2024
Closest in time.
A comprehensive survey on process-oriented automatic text summarization with exploration of llm-based methods
Hanlei Jin, Yang Zhang, Dan Meng, Jun Wang, and Jinghua Tan. 2024 · 2024
Closest in time.
LLM Inference Series: 4. KV caching, a deeper look
Pierre Lienhart. 2024 · 2024
Closest in time.
Andes: Defining and Enhancing Quality-of-Experience in LLM-Based Text Streaming Services
Jiachen Liu, Zhiyu Wu, Jae-Won Chung, Fan Lai, Myungjin Lee, and Mosharaf Chowdhury. 2024 · 2024
Closest in time.
Nvidia Multi-instance GPU
NVIDIA. 2024a · 2024
Closest in time.
TensorRT-LLM
NVIDIA. 2024b · 2024
Closest in time.
OpenAI - Finetuning
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023 · 2023
Cited alongside, same era.
Fast Inference from Transformers via Speculative Decoding. In Proceedings of the 40th International Conference on Machine Learning (ICML 2023) . PMLR, PMLR, Honolulu, HI, 19274–19286
Yaniv Leviathan, Matan Kalman, and Yossi Matias. 2023 · 2023
Cited alongside, same era.
AlpaServe: Statistical Multiplexing with Model Parallelism for Deep Learning Serving. In Proceedings of the 17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 2023) . USENIX, Boston, MA, 663–679
Zhuohan Li, Lianmin Zheng, Yinmin Zhong, Vincent Liu, Ying Sheng, Xin Jin, Yanping Huang, Zhifeng Chen, Hao Zhang, Joseph E Gonzalez, et al · 2023
Cited alongside, same era.
DejaVu: Contextual Sparsity for Efficient LLMs at Inference Time. In Proceedings of the 40th International Conference on Machine Learning (ICML 2023) . PMLR, PMLR, Honolulu, HI, 22137–22176
Zichang Liu, Jue Wang, Tri Dao, Tianyi Zhou, Binhang Yuan, Zhao Song, Anshumali Shrivastava, Ce Zhang, Yuandong Tian, Christopher Re, et al · 2023
Cited alongside, same era.
SpotServe: Serving generative large language models on preemptible instances
Xupeng Miao, Chunan Shi, Jiangfei Duan, Xiaoli Xi, Dahua Lin, Bin Cui, and Zhihao Jia. 2023 · 2023
Cited alongside, same era.
Splitwise: Efficient generative LLM inference using phase splitting
Pratyush Patel, Esha Choukse, Chaojie Zhang, Íñigo Goiri, Aashaka Shah, Saeed Maleki, and Ricardo Bianchini. 2023 · 2023
Cited alongside, same era.
Code llama: Open foundation models for code
Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, et al · 2023
Cited alongside, same era.
FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU. In Proceedings of the 40th International Conference on Machine Learning (ICML 2023) . PMLR, PMLR, Honolulu, HI, 31094–31116
Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher Ré, Ion Stoica, and Ce Zhang. 2023b · 2023
Cited alongside, same era.
OpenAI. 2024 · 2024
Closest in time.
Efficient Interactive LLM Serving with Proxy Model-based Sequence Length Prediction. In The 5th International Workshop on Cloud Intelligence / AIOps at ASPLOS 2024 , Vol. 5. Association for Computing Machinery, San Diego, CA, USA, 1–7
Haoran Qiu, Weichao Mao, Archit Patke, Shengkun Cui, Saurabh Jha, Chen Wang, Hubertus Franke, Zbigniew T. Kalbarczyk, Tamer Başar, and Ravishankar K. Iyer. 2024a · 2024
Closest in time.
Power-aware Deep Learning Model Serving with μ \mu -Serve. In Proceedings of the 2024 USENIX Annual Technical Conference (USENIX ATC 2024) . USENIX, Santa Clara, CA, 75–93
Haoran Qiu, Weichao Mao, Archit Patke, Shengkun Cui, Saurabh Jha, Chen Wang, Hubertus Franke, Zbigniew T Kalbarczyk, Tamer Başar, and Ravishankar K Iyer. 2024b · 2024
Closest in time.
RabbitMQ
RabbitMQ. 2024 · 2024
Closest in time.
Lab: Large-scale alignment for chatbots
Shivchander Sudalairaj, Abhishek Bhandwaldar, Aldo Pareja, Kai Xu, David D Cox, and Akash Srivastava. 2024 · 2024
Closest in time.
Llumnix: Dynamic Scheduling for Large Language Model Serving
Biao Sun, Ziming Huang, Hanyu Zhao, Wencong Xiao, Xinyi Zhang, Yong Li, and Wei Lin. 2024 · 2024
Closest in time.
ShareGPT Dataset
Vicuna team. 2024 · 2024
Closest in time.
Towards Efficient and Reliable LLM Serving: A Real-World Workload Study
Yuxin Wang, Yuhan Chen, Zeyu Li, Zhenheng Tang, Rui Guo, Xin Wang, Qiang Wang, Amelie Chi Zhou, and Xiaowen Chu. 2024 · 2024
Closest in time.
The Shift from Models to Compound AI Systems
Matei Zaharia, Omar Khattab, Lingjiao Chen, Jared Quincy Davis, Heather Miller, Chris Potts, James Zou, Michael Carbin, Jonathan Frankle, Naveen Rao, and Ali Ghodsi. 2024 · 2024
Closest in time.
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang. 2024 · 2024
Closest in time.
RelayAttention for Efficient Large Language Model Serving with Long System Prompts
Lei Zhu, Xinjiang Wang, Wayne Zhang, and Rynson W. H. Lau. 2024 · 2024
Closest in time.