Fetching the paper…
Reading the bibliography…
The widespread adoption of Large Language Models (LLMs) has enabled diverse applications with very different latency requirements.
Scheduling Algorithms for Multiprogramming in a Hard-Real-Time Environment
C. L. Liu and James W. Layland. 1973 · 1973
Earlier work this paper cites.
The Rate Monotonic Scheduling Algorithm: Exact Characterization and Average Case Behavior. In Proceedings of the Real-Time Systems Symposium - 1989, Santa Monica, California, USA, December 1989 . IEEE Computer Society, 166–171
John P. Lehoczky, Lui Sha, and Y. Ding. 1989 · 1989
Earlier work this paper cites.
Scheduling algorithms for distributed Web servers. In Proceedings of 17th International Conference on Distributed Computing Systems . 169–176
M. Colajanni, P.S. Yu, and D.M. Dias. 1997 · 1997
Earlier work this paper cites.
Web servers under overload: How scheduling can help
Bianca Schroeder and Mor Harchol-Balter. 2006 · 2006
Earlier work this paper cites.
BlinkDB: queries with bounded errors and bounded response times on very large data. In Eighth Eurosys Conference 2013, EuroSys ’13, Prague, Czech Republic, April 14-17, 2013 , Zdenek Hanzálek, Hermann Härtig, Miguel Castro, and M. Frans Kaashoek (Eds.). ACM, 29–42
Sameer Agarwal, Barzan Mozafari, Aurojit Panda, Henry Milner, Samuel Madden, and Ion Stoica. 2013 · 2013
Earlier work this paper cites.
Paragon: QoS-aware scheduling for heterogeneous datacenters. In Architectural Support for Programming Languages and Operating Systems, ASPLOS 2013, Houston, TX, USA, March 16-20, 2013 , Vivek Sarkar and Rastislav Bodík (Eds.). ACM, 77–88
Christina Delimitrou and Christos Kozyrakis. 2013 · 2013
Earlier work this paper cites.
Quasar: resource-efficient and QoS-aware cluster management. In Architectural Support for Programming Languages and Operating Systems, ASPLOS 2014, Salt Lake City, UT, USA, March 1-5, 2014 , Rajeev Balasubramonian, Al Davis, and Sarita V. Adve (Eds.). ACM, 127–144
Christina Delimitrou and Christos Kozyrakis. 2014 · 2014
Earlier work this paper cites.
Handling overload. Site Reliability Engineering: How Google runs production systems
Alejandro Forero Cuervo. 2017a · 2017
Earlier work this paper cites.
Load balancing in the datacenter. Site Reliability Engineering: How Google runs production systems
Alejandro Forero Cuervo. 2017b · 2017
Earlier work this paper cites.
Addressing cascading failures. Site Reliability Engineering: How Google runs production systems
Mike Ulrich. 2017 · 2017
Cited alongside, same era.
Attention is All you Need. In Advances in Neural Information Processing Systems , I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.), Vol. 30. Curran Associates, Inc
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Taiji: managing global user traffic for large-scale internet services at the edge
David Chou, Tianyin Xu, Kaushik Veeraraghavan, Andrew Newell, Sonia Margulis, Lin Xiao, Pol Mauri Ruiz, Justin Meza, Kiryong Ha, Shruti Padmanabha, Kevin Cole, and Dmitri Perelman. 2019 · 2019
Cited alongside, same era.
Serving { \{ DNNs } \} like Clockwork: Performance Predictability from the Bottom Up. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20) . 443–462
Arpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao, Antoine Kaufmann, Ymir Vigfusson, and Jonathan Mace. 2020 · 2020
Cited alongside, same era.
Etalon: Holistic Performance Evaluation Framework for LLM Inference Systems
Amey Agrawal, Anmol Agarwal, Nitin Kedia, Jayashree Mohan, Souvik Kundu, Nipun Kwatra, Ramachandran Ramjee, and Alexey Tumanov. 2024 · 2024
Later among the works it cites.
Vidur: A Large-Scale Simulation Framework For LLM Inference
Amey Agrawal, Nitin Kedia, Jayashree Mohan, Ashish Panwar, Nipun Kwatra, Bhargav S Gulavani, Ramachandran Ramjee, and Alexey Tumanov. 2024a · 2024
Later among the works it cites.
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
Jovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Josep Torrellas, and Esha Choukse. 2024 · 2024
Later among the works it cites.
[RFC] Upstream Chunked Prefill #3130
vLLM Project. 2024 · 2024
Later among the works it cites.
OpenChat: Advancing Open-source Language Models with Mixed-Quality Data
Guan Wang, Sijie Cheng, Xianyuan Zhan, Xiangang Li, Sen Song, and Yang Liu. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
{ \{ INFaaS } \} : Automated Model-less Inference Serving. In 2021 USENIX Annual Technical Conference (USENIX ATC 21) . 397–411
Francisco Romero, Qian Li, Neeraja J Yadwadkar, and Christos Kozyrakis. 2021 · 2021
Cited alongside, same era.
Orca: A Distributed Serving System for Transformer-Based Generative Models. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22) . USENIX Association, Carlsbad, CA, 521–538
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun. 2022 · 2022
Cited alongside, same era.
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Amey Agrawal, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S. Gulavani, and Ramachandran Ramjee. 2023 · 2023
Cited alongside, same era.
Efficient Memory Management for Large Language Model Serving with PagedAttention (SOSP ’23) . Association for Computing Machinery, New York, NY, USA, 611–626
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023 · 2023
Cited alongside, same era.
Admission Control with Response Time Objectives for Low-latency Online Data Systems
Hao Xu and Juan A. Colmenares. 2023 · 2023
Cited alongside, same era.
Taming { \{ Throughput-Latency } \} tradeoff in { \{ LLM } \} inference with { \{ Sarathi-Serve } \} . In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) . 117–134
Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav Gulavani, Alexey Tumanov, and Ramachandran Ramjee. 2024b
Cited in the paper.
Later among the works it cites.
SGLang: Efficient Execution of Structured Language Model Programs
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng. 2024 · 2024
Later among the works it cites.
Serving Models, Fast and Slow: Optimizing Heterogeneous LLM Inferencing Workloads at Scale
Shashwat Jaiswal, Kunal Jain, Yogesh Simmhan, Anjaly Parayil, Ankur Mallick, Rujia Wang, Renee St Amant, Chetan Bansal, Victor Rühle, Anoop Kulkarni, et al · 2025
Closest in time.
vAttention: Dynamic Memory Management for Serving LLMs without PagedAttention
Ramya Prabhu, Ajay Nayak, Jayashree Mohan, Ramachandran Ramjee, and Ashish Panwar. 2025 · 2025
Closest in time.