Fetching the paper…
Reading the bibliography…
Large language models now serve millions of users daily, with providers incurring costs exceeding $700,000 per day.
On queues in heavy traffic
J. F. C. Kingman · 1962
Earlier work this paper cites.
On the maximum and minimum of partial sums of random variables
Peggy Strait · 1974
Earlier work this paper cites.
Strong approximations for markovian service networks
Avi Mandelbaum, William A Massey, and Martin I Reiman · 1998
Earlier work this paper cites.
Discrete-review policies for scheduling stochastic networks: Trajectory tracking and fluid-scale asymptotic optimality
Constantinos Maglaras · 2000
Earlier work this paper cites.
Optimal control of queueing networks: An approach via fluid models
Nicole Bäuerle · 2002
Earlier work this paper cites.
Applied probability and queues , volume 2
Søren Asmussen · 2003
Earlier work this paper cites.
The adwords problem: online keyword matching with budgeted bidders under random permutations
Nikhil R Devanur and Thomas P Hayes · 2009
Earlier work this paper cites.
Optimal online assignment with forecasts
Erik Vee, Sergei Vassilvitskii, and Jayavel Shanmugasundaram · 2010
Earlier work this paper cites.
A network of time-varying many-server fluid queues with customer abandonment
Yunan Liu and Ward Whitt · 2011
Earlier work this paper cites.
Online batch scheduling for flow objectives
Sungjin Im and Benjamin Moseley · 2013
Earlier work this paper cites.
Efficient online scheduling for deadline-sensitive jobs
Brendan Lucier, Ishai Menache, Joseph Naor, and Jonathan Yaniv · 2013
Earlier work this paper cites.
The sample complexity of revenue maximization
Richard Cole and Tim Roughgarden · 2014
Earlier work this paper cites.
Online unbounded batch scheduling on parallel machines with delivery times
Peihai Liu and Xiwen Lu · 2015
Earlier work this paper cites.
The power of optimization from samples
Eric Balkanski, Aviad Rubinstein, and Yaron Singer · 2016
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2019
Earlier work this paper cites.
Probability: theory and examples , volume 49
Rick Durrett · 2019
Earlier work this paper cites.
Estimation of energy consumption in machine learning
Eva García-Martín, Crefeda Faviola Rodrigues, Graham Riley, and Håkan Grahn · 2019
Earlier work this paper cites.
Energy and policy considerations for deep learning in nlp
Emma Strubell, Ananya Ganesh, and Andrew McCallum · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Cited alongside, same era.
Online scheduling via learned weights
Silvio Lattanzi, Thomas Lavastida, Benjamin Moseley, and Sergei Vassilvitskii · 2020
Cited alongside, same era.
Online batch scheduling of simple linear deteriorating jobs with incompatible families
Wenhua Li, Libo Wang, Xing Chai, and Hang Yuan · 2020
Cited alongside, same era.
Compute and energy consumption trends in deep learning inference
Radosvet Desislavov, Silverio Martínez-Fernández, and Xavier Franch · 2021
Cited alongside, same era.
Carbon emissions and large neural network training
David Patterson, Joseph Gonzalez, Quoc Le, Chen Liang, Lluis-Miquel Munguia, Daniel Rothchild, David So, Maud Texier, and Jeff Dean · 2021
Efficient llm scheduling by learning to rank
Yichao Fu, Siqi Zhu, Runlong Su, Aurick Qiao, Ion Stoica, and Hao Zhang · 2024
Later among the works it cites.
Kvquant: Towards 10 million context length llm inference with kv cache quantization
Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Michael W Mahoney, Sophia Shao, Kurt Keutzer, and Amir Gholami · 2024
Later among the works it cites.
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, et al · 2024
Later among the works it cites.
Gear: An efficient kv cache compression recipe for near-lossless generative inference of llm
Hao Kang, Qingru Zhang, Souvik Kundu, Geonhwa Jeong, Zaoxing Liu, Tushar Krishna, and Tuo Zhao · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Deepspeed-inference: Enabling efficient inference of transformer models at unprecedented scale
Reza Yazdani Aminabadi, Samyam Rajbhandari, Ammar Ahmad Awan, Cheng Li, Du Li, Elton Zheng, Olatunji Ruwase, Shaden Smith, Minjia Zhang, Jeff Rasley, and Yuxiong He · 2022
Cited alongside, same era.
Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity
William Fedus, Barret Zoph, and Noam Shazeer · 2022
Cited alongside, same era.
Copilot, 2022
GitHub · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Cited alongside, same era.
Efficient transformers: A survey
Yi Tay, Mostafa Dehghani, Dara Bahri, and Donald Metzler · 2022
Cited alongside, same era.
Sustainable ai: Environmental implications, challenges and opportunities
Carole-Jean Wu, Ramya Raghavendra, Udit Gupta, Bilge Acun, Newsha Ardalani, Kiwan Maeng, Gloria Chang, Fiona Aga, Jinshi Huang, Charles Bai, et al · 2022
Cited alongside, same era.
Orca: A distributed serving system for { \{ Transformer-Based } \} generative models
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun · 2022
Cited alongside, same era.
OpenAI · 2024
Later among the works it cites.
Splitwise: Efficient generative llm inference using phase splitting
Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, and Ricardo Bianchini · 2024
Later among the works it cites.
Sglang: Fast serving framework for large language models and vision language models
SGLang Team · 2024
Later among the works it cites.
{ \{ DistServe } \} : Disaggregating prefill and decoding for goodput-optimized large language model serving
Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang · 2024
Later among the works it cites.
Claude 3.7 sonnet, 2025
Anthropic · 2025
Closest in time.
Optimal scheduling algorithms for llm inference: Theory and practice
Agrim Bari, Parikshit Hegde, and Gustavo de Veciana · 2025
Closest in time.
Adaptively robust llm inference optimization under prediction uncertainty
Zixi Chen, Yinyu Ye, and Zijie Zhou · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025
DeepSeek-AI · 2025
Closest in time.
Online scheduling for llm inference with kv cache constraints, 2025
Patrick Jaillet, Jiashuo Jiang, Konstantina Mellou, Marco Molinaro, Chara Podimata, and Zijie Zhou · 2025
Closest in time.
Oneiros: Kv cache optimization through parameter remapping for multi-tenant llm serving
Ruihao Li, Shagnik Pal, Vineeth Narayan Pullu, Prasoon Sinha, Jeeho Ryoo, Lizy K. John, and Neeraja J. Yadwadkar · 2025
Closest in time.
Throughput-optimal scheduling algorithms for llm inference and ai agents, 2025b
Yueying Li, Jim Dai, and Tianyi Peng · 2025
Closest in time.
Llm serving optimization with variable prefill and decode lengths
Meixuan Wang, Yinyu Ye, and Zijie Zhou · 2025
Closest in time.
Llama 2 model documentation
Hugging Face · 2026
Closest in time.
Nvidia a100 tensor core gpu specifications
NVIDIA · 2026
Closest in time.
Peeling The Onion’s Layers: Large Language Models’ Cost, Context, and Feature Breakdown
Dylan Patel · 2026
Closest in time.
Optimization and tuning
vLLM Team · 2026
Closest in time.