Fetching the paper…
Reading the bibliography…
Serving large language models (LLMs) for massive users is challenged by the significant memory footprint of the transient state, known as the key-value (KV) cache, which scales with sequence length and number of requests.
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, in NIPS , 2020, pp. 1877–1901
1901
Earlier work this paper cites.
R. M. Karp, Reducibility among Combinatorial Problems . Springer US, 1972, pp. 85–103. [Online]. Available: https://doi.org/10.1007/978-1-4684-2001-2_9
2001
Earlier work this paper cites.
D. Meisner, B. T. Gold, and T. F. Wenisch, “Powernap: eliminating server idle power,” SIGARCH Comput. Archit. News , vol. 37, no. 1, p. 205–216, mar 2009. [Online]. Available: https://doi.org/10.1145/2528521.1508269
2009
Earlier work this paper cites.
W. Song, Z. Xiao, Q. Chen, and H. Luo, “Adaptive resource provisioning for the cloud using online bin packing,” IEEE Transactions on Computers , vol. 63, no. 11, pp. 2647–2660, 2014
2014
Earlier work this paper cites.
S. Kamali and A. López-Ortiz, “Efficient online strategies for renting servers in the cloud,” in SOFSEM , 2015, pp. 277–288
2015
Earlier work this paper cites.
J. Fang, Y. Yu, C. Zhao, and J. Zhou, “Turbotransformers: an efficient gpu serving system for transformer models,” in PPoPP , 2021, p. 389–402
2021
Earlier work this paper cites.
G.-I. Yu, J. S. Jeong, G.-W. Kim, S. Kim, and B.-G. Chun, “Orca: A distributed serving system for Transformer-Based generative models,” in OSDI , 2022, pp. 521–538
2022
Earlier work this paper cites.
R. Y. Aminabadi, S. Rajbhandari, A. A. Awan, C. Li, D. Li, E. Zheng, O. Ruwase, S. Smith, M. Zhang, J. Rasley, and Y. He, “Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,” in SC , 2022
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
R. Pope, S. Douglas, A. Chowdhery, J. Devlin, J. Bradbury, J. Heek, K. Xiao, S. Agrawal, and J. Dean, “Efficiently scaling transformer inference,” in MLSys , 2023, pp. 606–624
2023
Earlier work this paper cites.
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica, “Efficient memory management for large language model serving with pagedattention,” in SOSP , 2023, p. 611–626
2023
Earlier work this paper cites.
Z. Zhang, Y. Sheng, T. Zhou, T. Chen, L. Zheng, R. Cai, Z. Song, Y. Tian, C. Re, C. Barrett, Z. Wang, and B. Chen, “H2o: Heavy-hitter oracle for efficient generative inference of large language models,” in NIPS , 2023
2023
Earlier work this paper cites.
Z. Liu, A. Desai, F. Liao, W. Wang, V. Xie, Z. Xu, A. Kyrillidis, and A. Shrivastava, “Scissorhands: Exploiting the persistence of importance hypothesis for LLM KV cache compression at test time,” in NIPS , 2023
2023
Earlier work this paper cites.
Y. Sheng, L. Zheng, B. Yuan, Z. Li, M. Ryabinin, B. Chen, P. Liang, C. Ré, I. Stoica, and C. Zhang, “Flexgen: high-throughput generative inference of large language models with a single gpu,” in ICML , 2023
2023
Earlier work this paper cites.
L. Zheng, W.-L. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, H. Zhang, J. E. Gonzalez, and I. Stoica, “Judging LLM-as-a-judge with MT-bench and chatbot arena,” in NIPS Datasets and Benchmarks Track , 2023
2023
Earlier work this paper cites.
X. Geng, A. Gudibande, H. Liu, E. Wallace, P. Abbeel, S. Levine, and D. Song, “Koala: A dialogue model for academic research,” Blog post, April 2023. [Online]. Available: https://bair.berkeley.edu/blog/2023/04/03/koala/
2023
Cited alongside, same era.
Z. Zheng, X. Ren, F. Xue, Y. Luo, X. Jiang, and Y. You, “Response length perception and sequence scheduling: An LLM-empowered LLM inference pipeline,” in NIPS , 2023
2023
Cited alongside, same era.
OpenAI, “Gpt-4 technical report,” 2024. [Online]. Available: https://arxiv.org/abs/2303.08774
2024
Cited alongside, same era.
Q. Hu, Z. Ye, Z. Wang, G. Wang, M. Zhang, Q. Chen, P. Sun, D. Lin, X. Wang, Y. Luo, Y. Wen, and T. Zhang, “Characterization of large language model development in the datacenter,” in NSDI , 2024, pp. 709–729
2024
Cited alongside, same era.
B. Gao, Z. He, P. Sharma, Q. Kang, D. Jevdjic, J. Deng, X. Yang, Z. Yu, and P. Zuo, “Cost-Efficient large language model serving for multi-turn conversations with CachedAttention,” in ATC , 2024, pp. 111–126
2024
Later among the works it cites.
B. Sun, Z. Huang, H. Zhao, W. Xiao, X. Zhang, Y. Li, and W. Lin, “Llumnix: Dynamic scheduling for large language model serving,” in OSDI , 2024, pp. 173–191
2024
Later among the works it cites.
Y. Fu, L. Xue, Y. Huang, A.-O. Brabete, D. Ustiugov, Y. Patel, and L. Mai, “Serverlessllm: Low-latency serverless inference for large language models,” in OSDI , 2024, pp. 135–153
2024
Later among the works it cites.
A. Agrawal, N. Kedia, A. Panwar, J. Mohan, N. Kwatra, B. Gulavani, A. Tumanov, and R. Ramjee, “Taming throughput-latency tradeoff in llm inference with sarathi-serve,” in OSDI , 2024, pp. 117–134
2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. Patel, E. Choukse, C. Zhang, A. Shah, I. Goiri, S. Maleki, and R. Bianchini, “Splitwise: Efficient generative llm inference using phase splitting,” in ISCA , June 2024
2024
Cited alongside, same era.
Z. Hong, J. Lin, S. Guo, S. Luo, W. Chen, R. Wattenhofer, and Y. Yu, “Optimus: Warming serverless ml inference via inter-function model transformation,” in EuroSys . Association for Computing Machinery, 2024, p. 1039–1053
2024
Cited alongside, same era.
J. Chen, W. Xu, Z. Hong, S. Guo, H. Wang, J. Zhang, and D. Zeng, “Otas: An elastic transformer serving system via token adaptation,” in INFOCOM , 2024, pp. 1–10
2024
Cited alongside, same era.
S. Ye, J. Du, L. Zeng, W. Ou, X. Chu, Y. Lu, and X. Chen, “Galaxy: A resource-efficient collaborative edge ai system for in-situ transformer inference,” in INFOCOM , 2024, pp. 1–10
2024
Cited alongside, same era.
Y. Chen, S. Qian, H. Tang, X. Lai, Z. Liu, S. Han, and J. Jia, “LongloRA: Efficient fine-tuning of long-context large language models,” in ICLR , 2024
2024
Cited alongside, same era.
M. Adnan, A. Arunkumar, G. Jain, P. Nair, I. Soloveychik, and P. Kamath, “Keyformer: Kv cache reduction through key tokens selection for efficient generative inference,” in MLSys , 2024, pp. 114–127
2024
Cited alongside, same era.
G. Xiao, Y. Tian, B. Chen, S. Han, and M. Lewis, “Efficient streaming language models with attention sinks,” in ICLR , 2024
2024
Cited alongside, same era.
2024
Cited alongside, same era.
Y. Zhong, S. Liu, J. Chen, J. Hu, Y. Zhu, X. Liu, and X. Jin, “Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving,” in OSDI , 2024, pp. 193–210
2024
Later among the works it cites.
2024
Later among the works it cites.
B. Wu, R. Zhu, Z. Zhang, P. Sun, X. Liu, and X. Jin, “dLoRA: Dynamically orchestrating requests and adapters for LoRA LLM serving,” in OSDI , 2024, pp. 911–927
2024
Later among the works it cites.
L. Zheng, W.-L. Chiang, Y. Sheng, T. Li, S. Zhuang, Z. Wu, Y. Zhuang, Z. Li, Z. Lin, E. Xing, J. E. Gonzalez, I. Stoica, and H. Zhang, “LMSYS-chat-1m: A large-scale real-world LLM conversation dataset,” in ICLR , 2024
2024
Later among the works it cites.
W. Zhao, X. Ren, J. Hessel, C. Cardie, Y. Choi, and Y. Deng, “Wildchat: 1m chatGPT interaction logs in the wild,” in ICLR , 2024
2024
Later among the works it cites.
vLLM, “Easy, fast, and cheap llm serving for everyone,” 2024. [Online]. Available: https://github.com/vllm-project/vllm
2024
Later among the works it cites.
Ray, “Ray: a unified framework for scaling ai and python applications,” 2024. [Online]. Available: https://github.com/ray-project/ray
2024
Later among the works it cites.
Meta, “Gloo: Collective communications library with various primitives for multi-machine training,” 2024. [Online]. Available: https://github.com/facebookincubator/gloo
2024
Later among the works it cites.
OpenAI, “Openai platform document,” 2024. [Online]. Available: https://platform.openai.com/docs/models
2024
Later among the works it cites.
Anthropic, “Anthropic platform document,” 2024. [Online]. Available: https://docs.anthropic.com/en/docs/about-claude/models
2024
Later among the works it cites.