Fetching the paper…
Reading the bibliography…
Serving numerous users and requests concurrently requires good fairness in Large Language Models (LLMs) serving system.
A simple hardware buddy system memory allocator
Von Puttkamer, E · 1975
Earlier work this paper cites.
Frage: Frequency-agnostic word representation
Gong, C., He, D., Tan, X., Qin, T., Wang, L., and Liu, T.-Y · 2018
Earlier work this paper cites.
Self-attention with relative position representations
Shaw, P., Uszkoreit, J., and Vaswani, A · 2018
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Gao, J., Choi, Y., et al · 2020
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Earlier work this paper cites.
Larimi, S. S. N., Salami, B., Unsal, O. S., Kestelman, A. C., Sarbazi-Azad, H., and Mutlu, O · 2020
Earlier work this paper cites.
Deepspeed inference: Enabling efficient inference of transformer models at unprecedented scale, 2022
Aminabadi, R. Y., Rajbhandari, S., Zhang, M., Awan, A. A., Li, C., Li, D., Zheng, E., Rasley, J., Smith, S., Ruwase, O., and He, Y · 2022
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Xia, F., Chi, E., Le, Q. V., Zhou, D., et al · 2022
Earlier work this paper cites.
Sarathi: Efficient llm inference by piggybacking decodes with chunked prefills
Agrawal, A., Panwar, A., Mohan, J., Kwatra, N., Gulavani, B. S., and Ramjee, R · 2023
Earlier work this paper cites.
Model tells you what to discard: Adaptive kv cache compression for llms
Ge, S., Zhang, Y., Liu, L., Zhang, M., Han, J., and Gao, J · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J., Zhang, H., and Stoica, I · 2023
Earlier work this paper cites.
Codegen: An open large language model for code with multi-turn program synthesis, 2023
Nijkamp, E., Pang, B., Hayashi, H., Tu, L., Wang, H., Zhou, Y., Savarese, S., and Xiong, C · 2023
Cited alongside, same era.
Flexgen: High-throughput generative inference of large language models with a single gpu
Sheng, Y., Zheng, L., Yuan, B., Li, Z., Ryabinin, M., Chen, B., Liang, P., Ré, C., Stoica, I., and Zhang, C · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., et al · 2023
Cited alongside, same era.
Fast distributed inference serving for large language models
Wu, B., Zhong, Y., Zhang, Z., Huang, G., Liu, X., and Jin, X · 2023
Cited alongside, same era.
Efficiently programming large language models using sglang
Lightllm: A lightweight framework for large language model inference
ModelTC · 2024
Closest in time.
Nvidia tensorrt-llm
NVIDIA · 2024
Closest in time.
One queue is all you need: Resolving head-of-line blocking in large language model serving
Patke, A., Reddy, D., Jha, S., Qiu, H., Pinto, C., Cui, S., Narayanaswami, C., Kalbarczyk, Z., and Iyer, R · 2024
Closest in time.
Mooncake: Kimi’s kvcache-centric architecture for llm serving
Qin, R., Li, Z., He, W., Zhang, M., Wu, Y., Zheng, W., and Xu, X · 2024
Closest in time.
Sharegpt: Share your wildest chatgpt conversations with one click
ShareGPT · 2024
Closest in time.
Fairness in serving large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Zheng, L., Yin, L., Xie, Z., Huang, J., Sun, C., Hao Yu, C., Cao, S., Kozyrakis, C., Stoica, I., Gonzalez, J. E., et al · 2023
Cited alongside, same era.
Taming throughput-latency tradeoff in llm inference with sarathi-serve
Agrawal, A., Kedia, N., Panwar, A., Mohan, J., Kwatra, N., Gulavani, B. S., Tumanov, A., and Ramjee, R · 2024
Cited alongside, same era.
Gao, B., He, Z., Sharma, P., Kang, Q., Jevdjic, D., Deng, J., Yang, X., Yu, Z., and Zuo, P · 2024
Cited alongside, same era.
Hugging face large language models (llms)
Hugging Face · 2024
Cited alongside, same era.
Generating images with multimodal language models
Koh, J. Y., Fried, D., and Salakhutdinov, R. R · 2024
Cited alongside, same era.
Cachegen: Kv cache compression and streaming for fast large language model serving
Liu, Y., Li, H., Cheng, Y., Ray, S., Huang, Y., Zhang, Q., Du, K., Yao, J., Lu, S., Ananthanarayanan, G., et al · 2024
Cited alongside, same era.
Large language models: A survey
Minaee, S., Mikolov, T., Nikzad, N., Chenaghlu, M., Socher, R., Amatriain, X., and Gao, J · 2024
Cited alongside, same era.
Andes: Defining and enhancing quality-of-experience in llm-based text streaming services, 2024a
Liu, J., Wu, Z., Chung, J.-W., Lai, F., Lee, M., and Chowdhury, M
Cited in the paper.
Sheng, Y., Cao, S., Li, D., Zhu, B., Li, Z., Zhuo, D., Gonzalez, J. E., and Stoica, I · 2024
Closest in time.
Llumnix: Dynamic scheduling for large language model serving, 2024
Sun, B., Huang, Z., Zhao, H., Xiao, W., Zhang, X., Li, Y., and Lin, W · 2024
Closest in time.
Yang, A., Yang, B., Hui, B., Zheng, B., Yu, B., Zhou, C., Li, C., Li, C., Liu, D., Huang, F., et al · 2024
Closest in time.
Llm as a system service on mobile devices
Yin, W., Xu, M., Li, Y., and Liu, X · 2024
Closest in time.
Zhong, Y., Liu, S., Chen, J., Hu, J., Zhu, Y., Liu, X., Jin, X., and Zhang, H · 2024
Closest in time.
Multilingual machine translation with large language models: Empirical results and analysis, 2024
Zhu, W., Liu, H., Dong, Q., Xu, J., Huang, S., Kong, L., Chen, J., and Li, L · 2024
Closest in time.