Fetching the paper…
Reading the bibliography…
Efficient scheduling is crucial for interactive Large Language Model (LLM) applications, where low request completion time directly impacts user engagement.
Soap bubbles: Robust scheduling under adversarial noise
Ziv Scully and Mor Harchol-Balter · 2018
Earlier work this paper cites.
Designing and interpreting probes with control tasks
John Hewitt and Percy Liang · 2019
Earlier work this paper cites.
A structural probe for finding syntax in word representations
John Hewitt and Christopher D Manning · 2019
Earlier work this paper cites.
Shinjuku: Preemptive scheduling for { \{ μ \mu second-scale } \} tail latency
Kostis Kaffes, Timothy Chong, Jack Tigar Humphries, Adam Belay, David Mazières, and Christos Kozyrakis · 2019
Earlier work this paper cites.
Scheduling with predictions and the price of misprediction
Michael Mitzenmacher · 2019
Earlier work this paper cites.
Distilbert, a distilled version of bert: Smaller, faster, cheaper and lighter
V Sanh · 2019
Earlier work this paper cites.
Probing classifiers: Promises, shortcomings, and advances
Yonatan Belinkov · 2022
Earlier work this paper cites.
Introducing chatgpt
OpenAI · 2022
Cited alongside, same era.
Orca: A distributed serving system for { \{ Transformer-Based } \} generative models
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun · 2022
Cited alongside, same era.
s 3 s^{3} : Increasing gpu utilization during generative inference for higher throughput
Yunho Jin, Chun-Feng Wu, David Brooks, and Gu-Yeon Wei · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica · 2023
Cited alongside, same era.
Bayesian filtering and smoothing , volume 17
Simo Särkkä and Lennart Svensson · 2023
Cited alongside, same era.
Stanford alpaca: An instruction-following llama model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B. Hashimoto · 2023
Fast distributed inference serving for large language models
Bingyang Wu, Yinmin Zhong, Zili Zhang, Gang Huang, Xuanzhe Liu, and Xin Jin · 2023
Later among the works it cites.
Enabling efficient batch serving for lmaas via generation length prediction
Ke Cheng, Wen Hu, Zhi Wang, Peng Du, Jianguo Li, and Sheng Zhang · 2024
Closest in time.
Power-aware deep learning model serving with { \{ μ \mu -Serve } \}
Haoran Qiu, Weichao Mao, Archit Patke, Shengkun Cui, Saurabh Jha, Chen Wang, Hubertus Franke, Zbigniew Kalbarczyk, Tamer Başar, and Ravishankar K Iyer · 2024
Closest in time.
Skippredict: When to invest in predictions for scheduling
Rana Shahout and Michael Mitzenmacher · 2024
Closest in time.
Dynamollm: Designing llm inference clusters for performance and energy efficiency
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Efficient interactive llm serving with proxy model-based sequence length prediction
Haoran Qiu, Weichao Mao, Archit Patke, Shengkun Cui, Saurabh Jha, Chen Wang, Hubertus Franke, Zbigniew T Kalbarczyk, Tamer Başar, and Ravishankar K Iyer
Cited in the paper.
Jovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Josep Torrellas, and Esha Choukse · 2024
Closest in time.
Response length perception and sequence scheduling: An llm-empowered llm inference pipeline
Zangwei Zheng, Xiaozhe Ren, Fuzhao Xue, Yang Luo, Xin Jiang, and Yang You · 2024
Closest in time.