Fetching the paper…
Reading the bibliography…
We study the problem of optimizing Large Language Model (LLM) inference scheduling to minimize total latency.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 1901
Earlier work this paper cites.
A review of machine scheduling: Complexity, algorithms and approximability
Bo Chen, Chris N Potts, and Gerhard J Woeginger. 1998 · 1998
Earlier work this paper cites.
Resource-constrained project scheduling: Notation, classification, models, and methods
Peter Brucker, Andreas Drexl, Rolf Möhring, Klaus Neumann, and Erwin Pesch. 1999 · 1999
Earlier work this paper cites.
Parallel machine scheduling with splitting jobs
Wenxun Xing and Jiawei Zhang. 2000 · 2000
Earlier work this paper cites.
An approximation algorithm for scheduling two parallel machines with capacity constraints
Heng Yang, Yinyu Ye, and Jiawei Zhang. 2003 · 2003
Earlier work this paper cites.
The price of robustness
Dimitris Bertsimas and Melvyn Sim. 2004 · 2004
Earlier work this paper cites.
Handbook on scheduling: from theory to applications
Jacek Blazewicz, Klaus H Ecker, Erwin Pesch, Gunter Schmidt, and Jan Weglarz. 2007 · 2007
Earlier work this paper cites.
Logistics scheduling with batching and transportation
Bo Chen and Chung-Yee Lee. 2008 · 2008
Earlier work this paper cites.
Competitive two-agent scheduling and its applications
Joseph Y-T Leung, Michael Pinedo, and Guohua Wan. 2010 · 2010
Earlier work this paper cites.
A unified framework for dynamic prediction market design
Shipra Agrawal, Erick Delage, Mark Peters, Zizhuo Wang, and Yinyu Ye. 2011 · 2011
Earlier work this paper cites.
Parallel machine selection and job scheduling to minimize machine cost and job tardiness
Dong Cao, Mingyuan Chen, and Guohua Wan. 2005 · 2012
Earlier work this paper cites.
Metaheuristics for solving a multimodal home-healthcare scheduling problem
Gerhard Hiermann, Matthias Prandtstetter, Andrea Rendl, Jakob Puchinger, and Günther R Raidl. 2015 · 2015
Earlier work this paper cites.
Appointment scheduling with limited distributional information
Ho-Yin Mak, Ying Rong, and Jiawei Zhang. 2015 · 2015
Earlier work this paper cites.
Secretary and online matching problems with machine learned advice
Antonios Antoniadis, Themis Gouleakis, Pieter Kleer, and Pavel Kolev. 2020 · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020 · 2020
Earlier work this paper cites.
Character AI
Character. 2021 · 2021
Earlier work this paper cites.
GitHub Copilot
GitHub. 2021 · 2021
Cited alongside, same era.
Competitive caching with machine learned advice
Thodoris Lykouris and Sergei Vassilvitskii. 2021 · 2021
Cited alongside, same era.
Online bipartite matching with advice: Tight robustness-consistency tradeoffs for the two-stage model
Billy Jin and Will Ma. 2022 · 2022
Cited alongside, same era.
Perplexity AI
Perplexity. 2022 · 2022
Cited alongside, same era.
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, et al · 2022
Cited alongside, same era.
Orca: A distributed serving system for { \{ Transformer-Based } \} generative models. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22) . 521–538
Robust scheduling with gflownets
David W Zhang, Corrado Rainone, Markus Peschl, and Roberto Bondesan. 2023 · 2023
Later among the works it cites.
LMSYS-Chat-1M: A large-scale real-world LLM conversation dataset
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zhuohan Li, Zi Lin, Eric Xing, et al · 2023
Later among the works it cites.
Response length perception and sequence scheduling: An llm-empowered llm inference pipeline
Zangwei Zheng, Xiaozhe Ren, Fuzhao Xue, Yang Luo, Xin Jiang, and Yang You. 2023b · 2023
Later among the works it cites.
Vidur: A Large-Scale Simulation Framework For LLM Inference
Amey Agrawal, Nitin Kedia, Jayashree Mohan, Ashish Panwar, Nipun Kwatra, Bhargav Gulavani, Ramachandran Ramjee, and Alexey Tumanov. 2024a · 2024
Later among the works it cites.
Taming throughput-latency tradeoff in LLM inference with Sarathi-Serve
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun. 2022 · 2022
Cited alongside, same era.
Advance admission scheduling via resource satisficing
Minglong Zhou, Melvyn Sim, and Shao-Wei Lam. 2022 · 2022
Cited alongside, same era.
Sarathi: Efficient LLM inference by piggybacking decodes with chunked prefills
Amey Agrawal, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S Gulavani, and Ramachandran Ramjee. 2023 · 2023
Cited alongside, same era.
Amazon CodeWhisperer
Amazon. 2023 · 2023
Cited alongside, same era.
Evaluating the feasibility of ChatGPT in healthcare: an analysis of multiple clinical and research scenarios
Marco Cascella, Jonathan Montomoli, Valentina Bellini, and Elena Bignami. 2023 · 2023
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2023
Cited alongside, same era.
Online resource allocation with convex-set machine-learned advice
Negin Golrezaei, Patrick Jaillet, and Zijie Zhou. 2023 · 2023
Cited alongside, same era.
Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S Gulavani, Alexey Tumanov, and Ramachandran Ramjee. 2024b · 2024
Later among the works it cites.
KVDirect: Distributed Disaggregated LLM Inference
Shiyang Chen, Rain Jiang, Dezhi Yu, Jinlai Xu, Mengyuan Chao, Fanlong Meng, Chenyu Jiang, Wei Xu, and Hang Liu. 2024 · 2024
Later among the works it cites.
Efficient LLM Scheduling by Learning to Rank
Yichao Fu, Siqi Zhu, Runlong Su, Aurick Qiao, Ion Stoica, and Hao Zhang. 2024 · 2024
Later among the works it cites.
Efficient interactive LLM serving with proxy model-based sequence length prediction
Haoran Qiu, Weichao Mao, Archit Patke, Shengkun Cui, Saurabh Jha, Chen Wang, Hubertus Franke, Zbigniew T Kalbarczyk, Tamer Başar, and Ravishankar K Iyer. 2024 · 2024
Later among the works it cites.
Don’t Stop Me Now: Embedding Based Scheduling for LLMs
Rana Shahout, Eran Malach, Chunwei Liu, Weifan Jiang, Minlan Yu, and Michael Mitzenmacher. 2024 · 2024
Later among the works it cites.
Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving
Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang. 2024 · 2024
Later among the works it cites.
Optimizing LLM Inference: Fluid-Guided Online Scheduling with Memory Constraints
Ruicheng Ao, Gan Luo, David Simchi-Levi, and Xinshang Wang. 2025 · 2025
Closest in time.
Orlm: A customizable framework in training large models for automated optimization modeling
Chenyu Huang, Zhengyang Tang, Shixi Hu, Ruoqing Jiang, Xin Zheng, Dongdong Ge, Benyou Wang, and Zizhuo Wang. 2025 · 2025
Closest in time.
Online Scheduling for LLM Inference with KV Cache Constraints
Patrick Jaillet, Jiashuo Jiang, Konstantina Mellou, Marco Molinaro, Chara Podimata, and Zijie Zhou. 2025 · 2025
Closest in time.
Throughput-optimal scheduling algorithms for llm inference and ai agents
Yueying Li, Jim Dai, and Tianyi Peng. 2025 · 2025
Closest in time.
LLM Serving Optimization with Variable Prefill and Decode Lengths
Meixuan Wang, Yinyu Ye, and Zijie Zhou. 2025 · 2025
Closest in time.