Fetching the paper…
Reading the bibliography…
KV cache accelerates LLM inference by avoiding redundant computation, at the expense of memory.
Few-to-many: Incremental parallelism for reducing tail latency in interactive services
Md E Haque, Yong Hun Eom, Yuxiong He, Sameh Elnikety, Ricardo Bianchini, and Kathryn S McKinley. 2015 · 2015
Earlier work this paper cites.
GPU computing pipeline inefficiencies and optimization opportunities in heterogeneous CPU-GPU processors. In 2015 IEEE International Symposium on Workload Characterization . IEEE, 87–97
Joel Hestness, Stephen W Keckler, and David A Wood. 2015 · 2015
Earlier work this paper cites.
Resource central: Understanding and predicting workloads for improved resource management in large cloud platforms. In Proceedings of the 26th Symposium on Operating Systems Principles . 153–167
Eli Cortez, Anand Bonde, Alexandre Muzio, Mark Russinovich, Marcus Fontoura, and Ricardo Bianchini. 2017 · 2017
Earlier work this paper cites.
Zygos: Achieving low tail latency for microsecond-scale networked tasks. In Proceedings of the 26th Symposium on Operating Systems Principles . 325–341
George Prekas, Marios Kogias, and Edouard Bugnion. 2017 · 2017
Earlier work this paper cites.
Amdahl’s law for tail latency
Christina Delimitrou and Christos Kozyrakis. 2018 · 2018
Earlier work this paper cites.
Leveraging power-performance relationship of energy-efficient modern DRAM devices
Sukhan Lee, Hyunyoon Cho, Young Hoon Son, Yuhwan Ro, Nam Sung Kim, and Jung Ho Ahn. 2018 · 2018
Earlier work this paper cites.
A performance & power comparison of modern high-speed dram architectures. In Proceedings of the International Symposium on Memory Systems . 341–353
Shang Li, Dhiraj Reddy, and Bruce Jacob. 2018 · 2018
Earlier work this paper cites.
Shinjuku: Preemptive Scheduling for { \{ μ \mu second-scale } \} Tail Latency. In 16th USENIX Symposium on Networked Systems Design and Implementation (NSDI 19) . 345–360
Kostis Kaffes, Timothy Chong, Jack Tigar Humphries, Adam Belay, David Mazières, and Christos Kozyrakis. 2019 · 2019
Earlier work this paper cites.
{ \{ PipeSwitch } \} : Fast pipelined context switching for deep learning applications. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20) . 499–514
Zhihao Bai, Zhen Zhang, Yibo Zhu, and Xin Jin. 2020 · 2020
Earlier work this paper cites.
Gslice: controlled spatial sharing of gpus for a scalable inference platform. In Proceedings of the 11th ACM Symposium on Cloud Computing . 492–506
Aditya Dhakal, Sameer G Kulkarni, and KK Ramakrishnan. 2020 · 2020
Earlier work this paper cites.
Serving { \{ DNNs } \} like clockwork: Performance predictability from the bottom up. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20) . 443–462
Arpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao, Antoine Kaufmann, Ymir Vigfusson, and Jonathan Mace. 2020 · 2020
Earlier work this paper cites.
Protean: { \{ VM } \} allocation service at scale. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20) . 845–861
Ori Hadary, Luke Marshall, Ishai Menache, Abhisek Pan, Esaias E Greeff, David Dion, Star Dorminey, Shailesh Joshi, Yang Chen, Mark Russinovich, et al · 2020
Earlier work this paper cites.
Swapadvisor: Pushing deep learning beyond the gpu memory limit via smart swapping. In Proceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems . 1341–1355
Chien-Chin Huang, Gu Jin, and Jinyang Li. 2020 · 2020
Earlier work this paper cites.
Capuchin: Tensor-based gpu memory management for deep learning. In Proceedings of the Twenty-Fifth International Conference on Architectural Support for Programming Languages and Operating Systems . 891–905
Xuan Peng, Xuanhua Shi, Hulin Dai, Hai Jin, Weiliang Ma, Qian Xiong, Fan Yang, and Xuehai Qian. 2020 · 2020
Earlier work this paper cites.
Zero: Memory optimizations toward training trillion parameter models. In SC20: International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 1–16
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He. 2020 · 2020
Earlier work this paper cites.
DeepSpeed: System Optimizations Enable Training Deep Learning Models with Over 100 Billion Parameters. In Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining (Virtual Event, CA, USA) (KDD ’20) . Association for Computing Machinery, New York, NY, USA, 3505–3506
Jeff Rasley, Samyam Rajbhandari, Olatunji Ruwase, and Yuxiong He. 2020 · 2020
Earlier work this paper cites.
Serverless in the wild: Characterizing and optimizing the serverless workload at a large cloud provider. In 2020 USENIX annual technical conference (USENIX ATC 20) . 205–218
Mohammad Shahrad, Rodrigo Fonseca, Inigo Goiri, Gohar Chaudhry, Paul Batum, Jason Cooke, Eduardo Laureano, Colby Tresness, Mark Russinovich, and Ricardo Bianchini. 2020 · 2020
Earlier work this paper cites.
{ \{ AntMan } \} : Dynamic scaling on { \{ GPU } \} clusters for deep learning. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20) . 533–548
Wencong Xiao, Shiru Ren, Yong Li, Yang Zhang, Pengyang Hou, Zhi Li, Yihui Feng, Wei Lin, and Yangqing Jia. 2020 · 2020
Earlier work this paper cites.
Zico: Efficient { \{ GPU } \} memory sharing for concurrent { \{ DNN } \} training. In 2021 USENIX Annual Technical Conference (USENIX ATC 21) . 161–175
Gangmuk Lim, Jeongseob Ahn, Wencong Xiao, Youngjin Kwon, and Myeongjae Jeon. 2021 · 2021
Earlier work this paper cites.
Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning. In Proceedings of the international conference for high performance computing, networking, storage and analysis . 1–14
Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, and Yuxiong He. 2021 · 2021
Earlier work this paper cites.
{ \{ Zero-offload } \} : Democratizing { \{ billion-scale } \} model training. In 2021 USENIX Annual Technical Conference (USENIX ATC 21) . 551–564
Jie Ren, Samyam Rajbhandari, Reza Yazdani Aminabadi, Olatunji Ruwase, Shuangyan Yang, Minjia Zhang, Dong Li, and Yuxiong He. 2021 · 2021
Earlier work this paper cites.
{ \{ INFaaS } \} : Automated model-less inference serving. In 2021 USENIX Annual Technical Conference (USENIX ATC 21) . 397–411
Francisco Romero, Qian Li, Neeraja J Yadwadkar, and Christos Kozyrakis. 2021 · 2021
Earlier work this paper cites.
Serving heterogeneous machine learning models on { \{ Multi-GPU } \} servers with { \{ Spatio-Temporal } \} sharing. In 2022 USENIX Annual Technical Conference (USENIX ATC 22) . 199–216
Seungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park, Youngjin Kwon, and Jaehyuk Huh. 2022 · 2022
Earlier work this paper cites.
Melon: Breaking the memory wall for resource-efficient on-device machine learning. In Proceedings of the 20th Annual International Conference on Mobile Systems, Applications and Services . 450–463
Qipeng Wang, Mengwei Xu, Chao Jin, Xinran Dong, Jinliang Yuan, Xin Jin, Gang Huang, Yunxin Liu, and Xuanzhe Liu. 2022 · 2022
Earlier work this paper cites.
igniter: Interference-aware gpu resource provisioning for predictable dnn inference in the cloud
Fei Xu, Jianian Xu, Jiabin Chen, Li Chen, Ruitao Shang, Zhi Zhou, and Fangming Liu. 2022 · 2022
Cited alongside, same era.
Orca: A distributed serving system for { \{ Transformer-Based } \} generative models. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22) . 521–538
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun. 2022 · 2022
Cited alongside, same era.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al · 2022
Cited alongside, same era.
Chateval: Towards better llm-based evaluators through multi-agent debate
Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu. 2023 · 2023
Cited alongside, same era.
Junyou Li, Qin Zhang, Yangbin Yu, Qiang Fu, and Deheng Ye. 2024b · 2024
Later among the works it cites.
A survey on LLM-based multi-agent systems: workflow, infrastructure, and challenges
Xinyi Li, Sai Wang, Siqi Zeng, Yu Wu, and Yi Yang. 2024a · 2024
Later among the works it cites.
Cachegen: Kv cache compression and streaming for fast large language model serving. In Proceedings of the ACM SIGCOMM 2024 Conference . 38–56
Yuhan Liu, Hanchen Li, Yihua Cheng, Siddhant Ray, Yuyang Huang, Qizheng Zhang, Kuntai Du, Jiayi Yao, Shan Lu, Ganesh Ananthanarayanan, et al · 2024
Later among the works it cites.
Spotserve: Serving generative large language models on preemptible instances. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 . 1112–1127
Xupeng Miao, Chunan Shi, Jiangfei Duan, Xiaoli Xi, Dahua Lin, Bin Cui, and Zhihao Jia. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Efficient memory management for large language model serving with pagedattention. In Proceedings of the 29th Symposium on Operating Systems Principles . 611–626
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023 · 2023
Cited alongside, same era.
{ \{ AlpaServe } \} : Statistical Multiplexing with Model Parallelism for Deep Learning Serving. In 17th USENIX Symposium on Operating Systems Design and Implementation (OSDI 23) . 663–679
Zhuohan Li, Lianmin Zheng, Yinmin Zhong, Vincent Liu, Ying Sheng, Xin Jin, Yanping Huang, Zhifeng Chen, Hao Zhang, Joseph E Gonzalez, et al · 2023
Cited alongside, same era.
Efficiently scaling transformer inference
Reiner Pope, Sholto Douglas, Aakanksha Chowdhery, Jacob Devlin, James Bradbury, Jonathan Heek, Kefan Xiao, Shivani Agrawal, and Jeff Dean. 2023 · 2023
Cited alongside, same era.
Flexgen: High-throughput generative inference of large language models with a single gpu. In International Conference on Machine Learning . PMLR, 31094–31116
Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher Ré, Ion Stoica, and Ce Zhang. 2023 · 2023
Cited alongside, same era.
Multi-agent collaboration: Harnessing the power of intelligent llm agents
Yashar Talebirad and Amirhossein Nadiri. 2023 · 2023
Cited alongside, same era.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. 2023 · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Cited alongside, same era.
Scalable tail latency estimation for data center networks. In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23) . 685–702
Kevin Zhao, Prateesh Goyal, Mohammad Alizadeh, and Thomas E Anderson. 2023 · 2023
Cited alongside, same era.
An lpddr-based cxl-pnm platform for tco-efficient inference of transformer-based large language models. In 2024 IEEE International Symposium on High-Performance Computer Architecture (HPCA) . IEEE, 970–982
Sang-Soo Park, KyungSoo Kim, Jinin So, Jin Jung, Jonggeon Lee, Kyoungwan Woo, Nayeon Kim, Younghyun Lee, Hyungyo Kim, Yongsuk Kwon, et al · 2024
Later among the works it cites.
Splitwise: Efficient generative llm inference using phase splitting. In 2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA) . IEEE, 118–132
Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, and Ricardo Bianchini. 2024 · 2024
Later among the works it cites.
Queue management for slo-oriented large language model serving. In Proceedings of the 2024 ACM Symposium on Cloud Computing . 18–35
Archit Patke, Dhemath Reddy, Saurabh Jha, Haoran Qiu, Christian Pinto, Chandra Narayanaswami, Zbigniew Kalbarczyk, and Ravishankar Iyer. 2024 · 2024
Later among the works it cites.
SmartOClock: Workload-and risk-aware overclocking in the cloud. In 2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA) . IEEE, 437–451
Jovan Stojkovic, Pulkit A Misra, Íñigo Goiri, Sam Whitlock, Esha Choukse, Mayukh Das, Chetan Bansal, Jason Lee, Zoey Sun, Haoran Qiu, et al · 2024
Later among the works it cites.
Dynamollm: Designing llm inference clusters for performance and energy efficiency
Jovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Josep Torrellas, and Esha Choukse. 2024b · 2024
Later among the works it cites.
Orion: Interference-aware, fine-grained GPU sharing for ML applications. In Proceedings of the Nineteenth European Conference on Computer Systems . 1075–1092
Foteini Strati, Xianzhe Ma, and Ana Klimovic. 2024 · 2024
Later among the works it cites.
Llumnix: Dynamic scheduling for large language model serving. In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) . 173–191
Biao Sun, Ziming Huang, Hanyu Zhao, Wencong Xiao, Xinyi Zhang, Yong Li, and Wei Lin. 2024 · 2024
Later among the works it cites.
Fali Wang, Zhiwei Zhang, Xianren Zhang, Zongyu Wu, Tzuhao Mo, Qiuhao Lu, Wanjing Wang, Rui Li, Junjie Xu, Xianfeng Tang, et al · 2024
Later among the works it cites.
Improving GPU Multi-Tenancy Through Dynamic Multi-Instance GPU Reconfiguration
Tianyu Wang, Sheng Li, Bingyao Li, Yue Dai, Ao Li, Geng Yuan, Yufei Ding, Youtao Zhang, and Xulong Tang. 2024a · 2024
Later among the works it cites.
Pie: Pooling CPU Memory for LLM Inference
Yi Xu, Ziming Mao, Xiangxi Mo, Shu Liu, and Ion Stoica. 2024 · 2024
Later among the works it cites.
Training ultra long context language model with fully pipelined distributed transformer
Jinghan Yao, Sam Ade Jacobs, Masahiro Tanaka, Olatunji Ruwase, Aamir Shafi, Hari Subramoni, and Dhabaleswar K Panda. 2024 · 2024
Later among the works it cites.
Accelerating the training of large language models using efficient activation rematerialization and optimal hybrid parallelism. In 2024 USENIX Annual Technical Conference (USENIX ATC 24) . 545–561
Tailing Yuan, Yuliang Liu, Xucheng Ye, Shenglong Zhang, Jianchao Tan, Bin Chen, Chengru Song, and Di Zhang. 2024 · 2024
Later among the works it cites.
Neo: Saving gpu memory crisis with cpu offloading for online llm inference
Xuanlin Jiang, Yang Zhou, Shiyi Cao, Ion Stoica, and Minlan Yu. 2025 · 2025
Closest in time.
Jiin Kim, Byeongjun Shin, Jinha Chung, and Minsoo Rhu. 2025 · 2025
Closest in time.
FlexInfer: Flexible LLM Inference with CPU Computations. In Proceedings of the 8th MLSys Conference
Seonjin Na, Geonhwa Jeong, Byung Hoon Ahn, Aaron Jezghani, Jeffrey Young, Christopher J Hughes, Tushar Krishna, and Hyesoon Kim. 2025 · 2025
Closest in time.
vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention. In ASPLOS
Ramya Prabhu, Ajay Nayak, Jayashree Mohan, Ramachandran Ramjee, and Ashish Panwar. 2025 · 2025
Closest in time.
https://sharegpt.com/
ShareGPT Team. 2025 · 2025
Closest in time.
LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention
Shang Yang, Junxian Guo, Haotian Tang, Qinghao Hu, Guangxuan Xiao, Jiaming Tang, Yujun Lin, Zhijian Liu, Yao Lu, and Song Han. 2025 · 2025
Closest in time.
Prism: Unleashing GPU Sharing for Cost-Efficient Multi-LLM Serving
Shan Yu, Jiarong Xing, Yifan Qiao, Mingyuan Ma, Yangmin Li, Yang Wang, Shuo Yang, Zhiqiang Xie, Shiyi Cao, Ke Bao, et al · 2025
Closest in time.
Jenga: Effective Memory Management for Serving LLM with Heterogeneity
Chen Zhang, Kuntai Du, Shu Liu, Woosuk Kwon, Xiangxi Mo, Yufeng Wang, Xiaoxuan Liu, Kaichao You, Zhuohan Li, Mingsheng Long, et al · 2025
Closest in time.