Fetching the paper…
Reading the bibliography…
Large language model (LLM) serving demands low latency and high throughput, but high load variability makes it challenging to achieve high GPU utilization.
Clipper: A Low-Latency online prediction serving system
D. Crankshaw, X. Wang, G. Zhou, M. J. Franklin, J. E. Gonzalez, and I. Stoica · 2017
Earlier work this paper cites.
Zygos: Achieving low tail latency for microsecond-scale networked tasks
G. Prekas, M. Kogias, and E. Bugnion · 2017
Earlier work this paper cites.
Dureader: a chinese machine reading comprehension dataset from real-world applications, 2018
W. He, K. Liu, J. Liu, Y. Lyu, S. Zhao, X. Xiao, Y. Liu, Y. Wang, H. Wu, Q. She, X. Liu, T. Wu, and H. Wang · 2018
Earlier work this paper cites.
S. Narayan, S. B. Cohen, and M. Lapata · 2018
Earlier work this paper cites.
Parties: Qos-aware resource partitioning for multiple interactive services
S. Chen, C. Delimitrou, and J. F. Martínez · 2019
Earlier work this paper cites.
Multi-news: A large-scale multi-document summarization dataset and abstractive hierarchical model
A. Fabbri, I. Li, T. She, S. Li, and D. Radev · 2019
Earlier work this paper cites.
Shinjuku: Preemptive scheduling for usecond-scale tail latency
K. Kaffes, T. Chong, J. T. Humphries, A. Belay, D. Mazières, and C. Kozyrakis · 2019
Earlier work this paper cites.
fairseq: A fast, extensible toolkit for sequence modeling, 2019
M. Ott, S. Edunov, A. Baevski, A. Fan, S. Gross, N. Ng, D. Grangier, and M. Auli · 2019
Earlier work this paper cites.
Shenango: Achieving high CPU efficiency for latency-sensitive datacenter workloads
A. Ousterhout, J. Fried, J. Behrens, A. Belay, and H. Balakrishnan · 2019
Earlier work this paper cites.
The impact of gpu dvfs on the energy and performance of deep learning: an empirical study
Z. Tang, Y. Wang, Q. Wang, and X. Chu · 2019
Earlier work this paper cites.
PipeSwitch: Fast pipelined context switching for deep learning applications
Z. Bai, Z. Zhang, Y. Zhu, and X. Jin · 2020
Earlier work this paper cites.
Inferline: latency-aware provisioning and scaling for prediction serving pipelines
D. Crankshaw, G.-E. Sela, X. Mo, C. Zumar, I. Stoica, J. Gonzalez, and A. Tumanov · 2020
Earlier work this paper cites.
Caladan: Mitigating interference at microsecond timescales
J. Fried, Z. Ruan, A. Ousterhout, and A. Belay · 2020
Earlier work this paper cites.
Serving DNNs like clockwork: Performance predictability from the bottom up
A. Gujarati, R. Karimi, S. Alzayat, W. Hao, A. Kaufmann, Y. Vigfusson, and J. Mace · 2020
Earlier work this paper cites.
AntMan: Dynamic scaling on GPU clusters for deep learning
W. Xiao, S. Ren, Y. Li, Y. Zhang, P. Hou, Z. Li, Y. Feng, W. Lin, and Y. Jia · 2020
Earlier work this paper cites.
Deepspeed inference: Enabling efficient inference of transformer models at unprecedented scale, 2022
R. Y. Aminabadi, S. Rajbhandari, M. Zhang, A. A. Awan, C. Li, D. Li, E. Zheng, J. Rasley, S. Smith, O. Ruwase, and Y. He · 2022
Earlier work this paper cites.
Microsecond-scale preemption for concurrent GPU-accelerated DNN inferences
M. Han, H. Zhang, R. Chen, and H. Chen · 2022
Earlier work this paper cites.
Can foundation models wrangle your data?, 2022
A. Narayan, I. Chami, L. Orr, S. Arora, and C. Ré · 2022
Earlier work this paper cites.
Orca: A distributed serving system for Transformer-Based generative models
G.-I. Yu, J. S. Jeong, G.-W. Kim, S. Kim, and B.-G. Chun · 2022
Earlier work this paper cites.
Chatgpt sets record for fastest-growing user base
K. Hu · 2023
Earlier work this paper cites.
Llm-assisted code cleaning for training accurate code generators, 2023
N. Jain, T. Zhang, W.-L. Chiang, J. E. Gonzalez, K. Sen, and I. Stoica · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica · 2023
Earlier work this paper cites.
AlpaServe: Statistical multiplexing with model parallelism for deep learning serving
Z. Li, L. Zheng, Y. Zhong, V. Liu, Y. Sheng, X. Jin, Y. Huang, Z. Chen, H. Zhang, J. E. Gonzalez, and I. Stoica · 2023
Earlier work this paper cites.
Holistic evaluation of language models, 2023
P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, M. Yasunaga, Y. Zhang, D. Narayanan, Y. Wu, A. Kumar, B. Newman, B. Yuan, B. Yan, C. Zhang, C. Cosgrove, C. D. Manning, C. Ré, D. Acosta-Navas, D. A. Hudson, E. Zelikman, E. Durmus, F. Ladhak, F. Rong, H. Ren, H. Yao, J. Wang, K. Santhanam, L. Orr, L. Zheng, M. Yuksekgonul, M. Suzgun, N. Kim, N. Guha, N. Chatterji, O. Khattab, P. Henderson, Q. Huang, R. Chi, S. M. Xie, S. Santurkar, S. Ganguli, T. Hashimoto, T. Icard, T. Zhang, V. Chaudhary, W. Wang, X. Li, Y. Mai, Y. Zhang, and Y. Koreeda · 2023
Earlier work this paper cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Cited alongside, same era.
Occl: a deadlock-free library for gpu collective communication, 2023
L. Pan, J. Liu, J. Yuan, R. Zhang, P. Li, and Z. Xiao · 2023
Cited alongside, same era.
Llm has a performance problem inherent to its architecture: Latency
Proxet · 2023
Cited alongside, same era.
Flexgen: high-throughput generative inference of large language models with a single gpu
Y. Sheng, L. Zheng, B. Yuan, Z. Li, M. Ryabinin, B. Chen, P. Liang, C. Ré, I. Stoica, and C. Zhang · 2023
Cited alongside, same era.
Vcsum: A versatile chinese meeting summarization dataset, 2023
H. Wu, M. Zhan, H. Tan, Z. Hou, D. Liang, and L. Song · 2023
Cited alongside, same era.
SHEPHERD: Serving DNNs in the wild
H. Zhang, Y. Tang, A. Khandelwal, and I. Stoica · 2023
Cited alongside, same era.
Fastertransformer: Transformer related optimization, including bert, gpt
NVIDIA · 2024
Closest in time.
System Management Interface SMI
NVIDIA · 2024
Closest in time.
Tensorrt-llm: A tensorrt toolbox for optimized large language model inference
NVIDIA · 2024
Closest in time.
Triton Inference Server
NVIDIA · 2024
Closest in time.
Batch api
OpenAI · 2024
Closest in time.
Streaming api
OpenAI · 2024
Closest in time.
HotSpot Glossary of Terms: Safepoint
OpenJDK · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Efficiently programming large language models using sglang, 2023
L. Zheng, L. Yin, Z. Xie, J. Huang, C. Sun, C. H. Yu, S. Cao, C. Kozyrakis, I. Stoica, J. E. Gonzalez, C. Barrett, and Y. Sheng · 2023
Cited alongside, same era.
Taming Throughput-Latency tradeoff in LLM inference with Sarathi-Serve, July 2024
A. Agrawal, N. Kedia, A. Panwar, J. Mohan, N. Kwatra, B. Gulavani, A. Tumanov, and R. Ramjee · 2024
Cited alongside, same era.
Message batches
Anthropic · 2024
Cited alongside, same era.
Streaming messages
Anthropic · 2024
Cited alongside, same era.
Chatbot arena: An open platform for evaluating llms by human preference, 2024
W.-L. Chiang, L. Zheng, Y. Sheng, A. N. Angelopoulos, T. Li, D. Li, H. Zhang, B. Zhu, M. Jordan, J. E. Gonzalez, and I. Stoica · 2024
Cited alongside, same era.
DeepSeek-V3/R1 Inference System Overview
DeepSeek AI · 2024
Cited alongside, same era.
Splitwise: Efficient generative llm inference using phase splitting, 2024
P. Patel, E. Choukse, C. Zhang, A. Shah, Íñigo Goiri, S. Maleki, and R. Bianchini · 2024
Closest in time.
gc—Garbage Collector interface
Python Software Foundation · 2024
Closest in time.
Code llama: Open foundation models for code, 2024
B. Rozière, J. Gehring, F. Gloeckle, S. Sootla, I. Gat, X. E. Tan, Y. Adi, J. Liu, R. Sauvestre, T. Remez, J. Rapin, A. Kozhevnikov, I. Evtimov, J. Bitton, M. Bhatt, C. C. Ferrer, A. Grattafiori, W. Xiong, A. Défossez, J. Copet, F. Azhar, H. Touvron, L. Martin, N. Usunier, T. Scialom, and G. Synnaeve · 2024
Closest in time.
Fairness in serving large language models
Y. Sheng, S. Cao, D. Li, B. Zhu, Z. Li, D. Zhuo, J. E. Gonzalez, and I. Stoica · 2024
Closest in time.
Introducing the next generation of claude
The Claude Team · 2024
Closest in time.
Torchserve is a performant, flexible, and easy to use tool for serving pytorch models in production
The PyTorch Foundation · 2024
Closest in time.
Towards efficient and reliable llm serving: A real-world workload study, 2024
Y. Wang, Y. Chen, Z. Li, Z. Tang, R. Guo, X. Wang, Q. Wang, A. C. Zhou, and X. Chu · 2024
Closest in time.
A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Xia, X. Ren, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Wan, Y. Liu, Z. Cui, Z. Zhang, and Z. Qiu · 2024
Closest in time.
Fast and live model auto scaling with o(1) host caching, 2024
D. Zhang, H. Wang, Y. Liu, X. Wei, Y. Shan, R. Chen, and H. Chen · 2024
Closest in time.
Blendserve: Optimizing offline inference for auto-regressive large models with resource-aware batching, 2024
Y. Zhao, S. Yang, K. Zhu, L. Zheng, B. Kasikci, Y. Zhou, J. Xing, and I. Stoica · 2024
Closest in time.
DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving, July 2024
Y. Zhong, S. Liu, J. Chen, J. Hu, Y. Zhu, X. Liu, X. Jin, and H. Zhang · 2024
Closest in time.
Github copilot - write code faster
GitHub · 2025
Closest in time.
Conditional CUDA Graph Nodes
NVIDIA · 2025
Closest in time.
Flexllm: A system for co-serving large language model inference and parameter-efficient finetuning, 2025
G. Oliaro, X. Miao, X. Cheng, V. Kada, R. Gao, Y. Huang, R. Delacourt, A. Yang, Y. Wang, M. Wu, C. Unger, and Z. Jia · 2025
Closest in time.
Chatgpt: Conversational language model
OpenAI · 2025
Closest in time.
vLLM Piecewise CUDA Graph
vLLM team · 2025
Closest in time.
Skylb: A locality-aware cross-region load balancer for llm inference
T. Xia, Z. Mao, J. Kerney, E. J. Jackson, Z. Li, J. Xing, S. Shenker, and I. Stoica · 2025
Closest in time.