Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) show great capabilities in a wide range of applications, but serving them efficiently becomes increasingly challenging as requests (prompts) become more complex.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Longformer: The long-document transformer
Beltagy, I., Peters, M. E., and Cohan, A · 2020
Earlier work this paper cites.
Generation-augmented retrieval for open-domain question answering
Mao, Y., He, P., Liu, X., Shen, Y., Gao, J., Han, J., and Chen, W · 2021
Earlier work this paper cites.
A survey on retrieval-augmented text generation
Li, H., Su, Y., Cai, D., Wang, Y., and Liu, L · 2022
Earlier work this paper cites.
Orca: A distributed serving system for transformer-based generative models
Yu, G., Jeong, J. S., Kim, G., Kim, S., and Chun, B · 2022
Earlier work this paper cites.
Retrieval-augmented generation for large language models: A survey
Gao, Y., Xiong, Y., Gao, X., Jia, K., Pan, J., Bi, Y., Dai, Y., Sun, J., Guo, Q., Wang, M., and Wang, H · 2023
Earlier work this paper cites.
Predicting compilation resources for adaptive build in an industrial setting
Hu, J., Wang, C., Huang, H., Luo, H., Jin, Y., Deng, Y., and Xie, T · 2023
Earlier work this paper cites.
Mistral 7b
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de Las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with PagedAttention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
Earlier work this paper cites.
In-context retrieval-augmented language models
Ram, O., Levine, Y., Dalmedigos, I., Muhlgay, D., Shashua, A., Leyton-Brown, K., and Shoham, Y · 2023
Earlier work this paper cites.
H2O: heavy-hitter oracle for efficient generative inference of large language models
Zhang, Z., Sheng, Y., Zhou, T., Chen, T., Zheng, L., Cai, R., Song, Z., Tian, Y., Ré, C., Barrett, C. W., Wang, Z., and Chen, B · 2023
Cited alongside, same era.
LongBench: A bilingual, multitask benchmark for long context understanding
Bai, Y., Lv, X., Zhang, J., Lyu, H., Tang, J., Huang, Z., Du, Z., Liu, X., Zeng, A., Hou, L., Dong, Y., Tang, J., and Li, J · 2024
Cited alongside, same era.
ArkVale: Efficient generative LLM inference with recallable key-value eviction
Chen, R., Wang, Z., Cao, B., Wu, T., Zheng, S., Li, X., Wei, X., Yan, S., Li, M., and Liang, Y · 2024
Cited alongside, same era.
The Llama 3 herd of models
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., Goyal, A., Hartshorn, A., Yang, A., Mitra, A., Sravankumar, A., Korenev, A., Hinsvark, A., Rao, A., Zhang, A., Rodriguez, A., Gregerson, A., Spataru, A., Rozière, B., Biron, B., Tang, B., Chern, B., Caucheteux, C., Nayak, C., Bi, C., Marra, C., McConnell, C., Keller, C., Touret, C., Wu, C., Wong, C., Ferrer, C. C., Nikolaidis, C., Allonsius, D., Song, D., Pintz, D., Livshits, D., Esiobu, D., Choudhary, D., Mahajan, D., Garcia-Olano, D., Perino, D., Hupkes, D., Lakomkin, E., AlBadawy, E., Lobanova, E., Dinan, E., Smith, E. M., Radenovic, F., Zhang, F., Synnaeve, G., Lee, G., Anderson, G. L., Nail, G., Mialon, G., Pang, G., Cucurell, G., Nguyen, H., Korevaar, H., Xu, H., Touvron, H., Zarov, I., Ibarra, I. A., Kloumann, I. M., Misra, I., Evtimov, I., Copet, J., Lee, J., Geffert, J., Vranes, J., Park, J., Mahadeokar, J., Shah, J., van der Linde, J., Billock, J., Hong, J., Lee, J., Fu, J., Chi, J., Huang, J., Liu, J., Wang, J., Yu, J., Bitton, J., Spisak, J., Park, J., Rocca, J., Johnstun, J., Saxe, J., Jia, J., Alwala, K. V., Upasani, K., Plawiak, K., Li, K., Heafield, K., Stone, K., and et al · 2024
SLoRA: Scalable serving of thousands of lora adapters
Sheng, Y., Cao, S., Li, D., Hooper, C., Lee, N., Yang, S., Chou, C., Zhu, B., Zheng, L., Keutzer, K., Gonzalez, J., and Stoica, I · 2024
Closest in time.
QUEST: query-aware sparsity for efficient long-context LLM inference
Tang, J., Zhao, Y., Zhu, K., Xiao, G., Kasikci, B., and Han, S · 2024
Closest in time.
Efficient streaming language models with attention sinks
Xiao, G., Tian, Y., Chen, B., Han, S., and Lewis, M · 2024
Closest in time.
Yi: Open foundation models by 01.AI
Young, A., Chen, B., Li, C., Huang, C., Zhang, G., Zhang, G., Li, H., Zhu, J., Chen, J., Chang, J., Yu, K., Liu, P., Liu, Q., Yue, S., Yang, S., Yang, S., Yu, T., Xie, W., Huang, W., Hu, X., Ren, X., Niu, X., Nie, P., Xu, Y., Liu, Y., Wang, Y., Cai, Y., Gu, Z., Liu, Z., and Dai, Z · 2024
Closest in time.
SGLang: Efficient execution of structured language model programs
Zheng, L., Yin, L., Xie, Z., Sun, C., Huang, J., Yu, C. H., Cao, S., Kozyrakis, C., Stoica, I., Gonzalez, J. E., Barrett, C. W., and Sheng, Y · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Prompt cache: Modular attention reuse for low-latency inference
Gim, I., Chen, G., Lee, S., Sarda, N., Khandelwal, A., and Zhong, L · 2024
Cited alongside, same era.
Adaptive-RAG: Learning to adapt retrieval-augmented large language models through question complexity
Jeong, S., Baek, J., Cho, S., Hwang, S. J., and Park, J · 2024
Cited alongside, same era.
RAGCache: Efficient knowledge caching for retrieval-augmented generation
Jin, C., Zhang, Z., Jiang, X., Liu, F., Liu, X., Liu, X., and Jin, X · 2024
Cited alongside, same era.
CaraServe: CPU-assisted and rank-aware LoRA serving for generative LLM inference
Li, S., Lu, H., Wu, T., Yu, M., Weng, Q., Chen, X., Shan, Y., Yuan, B., and Wang, W · 2024
Cited alongside, same era.
Cachegen: KV cache compression and streaming for fast large language model serving
Liu, Y., Li, H., Cheng, Y., Ray, S., Huang, Y., Zhang, Q., Du, K., Yao, J., Lu, S., Ananthanarayanan, G., Maire, M., Hoffmann, H., Holtzman, A., and Jiang, J · 2024
Cited alongside, same era.
Splitwise: Efficient generative LLM inference using phase splitting
Patel, P., Choukse, E., Zhang, C., Shah, A., Goiri, Í., Maleki, S., and Bianchini, R · 2024
Cited alongside, same era.
https://github.com/gkamradt/LLMTest_NeedleInAHaystack
Needle In A Haystack
Cited in the paper.
https://gemini.google.com/ , a
Gemini
Cited in the paper.
DistServe: Disaggregating prefill and decoding for goodput-optimized LLM serving
Zhong, Y., Liu, S., Chen, J., Hu, J., Zhu, Y., Liu, X., Jin, X., and Zhang, H · 2024
Closest in time.
A survey on efficient inference for large language models
Zhou, Z., Ning, X., Hong, K., Fu, T., Xu, J., Li, S., Lou, Y., Wang, L., Yuan, Z., Li, X., Yan, S., Dai, G., Zhang, X., Dong, Y., and Wang, Y · 2024
Closest in time.
Mooncake: Trading more storage for less computation - A KVCache-centric architecture for serving LLM chatbot
Qin, R., Li, Z., He, W., Cui, J., Ren, F., Zhang, M., Wu, Y., Zheng, W., and Xu, X · 2025
Closest in time.
CacheBlend: Fast large language model serving for RAG with cached knowledge fusion
Yao, J., Li, H., Liu, Y., Ray, S., Cheng, Y., Zhang, Q., Du, K., Lu, S., and Jiang, J · 2025
Closest in time.
Stateful large language model serving with Pensieve
Yu, L., Lin, J., and Li, J · 2025
Closest in time.