Fetching the paper…
Reading the bibliography…
Serving Large Language Models (LLMs) efficiently in multi-region setups remains a challenge.
Consistent hashing and random trees: distributed caching protocols for relieving hot spots on the World Wide Web. In Proceedings of the Twenty-Ninth Annual ACM Symposium on Theory of Computing (El Paso, Texas, USA) (STOC ’97) . Association for Computing Machinery, New York, NY, USA, 654–663
David Karger, Eric Lehman, Tom Leighton, Rina Panigrahy, Matthew Levine, and Daniel Lewin. 1997 · 1997
Earlier work this paper cites.
Chord: a scalable peer-to-peer lookup protocol for internet applications
Ion Stoica, Robert Morris, David Liben-Nowell, David R Karger, M Frans Kaashoek, Frank Dabek, and Hari Balakrishnan. 2003 · 2003
Earlier work this paper cites.
Scalable work stealing. In Proceedings of the Conference on High Performance Computing Networking, Storage and Analysis . 1–11
James Dinan, D Brian Larkins, Ponnuswamy Sadayappan, Sriram Krishnamoorthy, and Jarek Nieplocha. 2009 · 2009
Earlier work this paper cites.
Sender initiated decentralized dynamic load balancing for multi cluster computational grid environment
Malarvizhi Nandagopal, Kandaswamy Gokulnath, and V Rhymend Uthariaraj. 2010 · 2010
Earlier work this paper cites.
Work stealing for interactive services to meet target latency. In Proceedings of the 21st ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming . 1–13
Jing Li, Kunal Agrawal, Sameh Elnikety, Yuxiong He, I-Ting Angelina Lee, Chenyang Lu, and Kathryn S McKinley. 2016 · 2016
Earlier work this paper cites.
ZygOS: Achieving Low Tail Latency for Microsecond-scale Networked Tasks. In Proceedings of the 26th ACM Symposium on Operating Systems Principles
Edouard Bugnion. 2017 · 2017
Earlier work this paper cites.
Vesper: Measuring Time-to-Interactivity for Web Pages. In 15th USENIX Symposium on Networked Systems Design and Implementation (NSDI 18) . USENIX Association, Renton, WA, 217–231
Ravi Netravali, Vikram Nathan, James Mickens, and Hari Balakrishnan. 2018 · 2018
Earlier work this paper cites.
Gandiva: Introspective cluster scheduling for deep learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18) . 595–610
Wencong Xiao, Romil Bhardwaj, Ramachandran Ramjee, Muthian Sivathanu, Nipun Kwatra, Zhenhua Han, Pratyush Patel, Xuan Peng, Hanyu Zhao, Quanlu Zhang, et al · 2018
Earlier work this paper cites.
Shenango: Achieving high CPU efficiency for latency-sensitive datacenter workloads. In 16th USENIX Symposium on Networked Systems Design and Implementation (NSDI 19) . 361–378
Amy Ousterhout, Joshua Fried, Jonathan Behrens, Adam Belay, and Hari Balakrishnan. 2019 · 2019
Earlier work this paper cites.
MArk: Exploiting cloud services for Cost-Effective,SLO-Aware machine learning inference serving. In 2019 USENIX Annual Technical Conference (USENIX ATC 19) . 1049–1062
Chengliang Zhang, Minchen Yu, Wei Wang, and Feng Yan. 2019 · 2019
Earlier work this paper cites.
Gslice: controlled spatial sharing of gpus for a scalable inference platform. In Proceedings of the 11th ACM Symposium on Cloud Computing . 492–506
Aditya Dhakal, Sameer G Kulkarni, and KK Ramakrishnan. 2020 · 2020
Earlier work this paper cites.
Caladan: Mitigating interference at microsecond timescales. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20) . 281–297
Joshua Fried, Zhenyuan Ruan, Amy Ousterhout, and Adam Belay. 2020 · 2020
Earlier work this paper cites.
User-level threading: Have your cake and eat it too
Martin Karsten and Saman Barghi. 2020 · 2020
Earlier work this paper cites.
AntMan: Dynamic scaling on GPU clusters for deep learning. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20) . 533–548
Wencong Xiao, Shiru Ren, Yong Li, Yang Zhang, Pengyang Hou, Zhi Li, Yihui Feng, Wei Lin, and Yangqing Jia. 2020 · 2020
Earlier work this paper cites.
Multi-model machine learning inference serving with gpu spatial partitioning
Seungbeom Choi, Sunho Lee, Yeonjae Kim, Jongse Park, Youngjin Kwon, and Jaehyuk Huh. 2021 · 2021
Earlier work this paper cites.
Training Verifiers to Solve Math Word Problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman. 2021 · 2021
Earlier work this paper cites.
Scalable Load Balancing in Networked Systems: A Survey of Recent Advances
Mark Van der Boor, Sem C. Borst, Johan S. H. Van Leeuwaarden, and Debankur Mukherjee. 2022 · 2022
Earlier work this paper cites.
Efficient scheduling policies for Microsecond-Scale tasks. In 19th USENIX Symposium on Networked Systems Design and Implementation (NSDI 22) . 1–18
Sarah McClure, Amy Ousterhout, Scott Shenker, and Sylvia Ratnasamy. 2022 · 2022
Earlier work this paper cites.
Orca: A Distributed Serving System for Transformer-Based Generative Models. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22) . USENIX Association, Carlsbad, CA, 521–538
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun. 2022 · 2022
Earlier work this paper cites.
Norman P. Jouppi, George Kurian, Sheng Li, Peter Ma, Rahul Nagarajan, Lifeng Nai, Nishant Patil, Suvinay Subramanian, Andy Swing, Brian Towles, Cliff Young, Xiang Zhou, Zongwei Zhou, and David Patterson. 2023 · 2023
Earlier work this paper cites.
Efficient Memory Management for Large Language Model Serving with PagedAttention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica. 2023 · 2023
Earlier work this paper cites.
Colti: Towards concurrent and co-located dnn training and inference. In Proceedings of the 32nd International Symposium on High-Performance Parallel and Distributed Computing . 309–310
Jaiaid Mobin, Avinash Maurya, and M Mustafa Rafique. 2023 · 2023
Earlier work this paper cites.
SkyPilot: An intercloud broker for sky computing. In 20th USENIX Symposium on Networked Systems Design and Implementation (NSDI 23) . 437–455
Zongheng Yang, Zhanghao Wu, Michael Luo, Wei-Lin Chiang, Romil Bhardwaj, Woosuk Kwon, Siyuan Zhuang, Frank Sifei Luan, Gautam Mittal, Scott Shenker, et al · 2023
Earlier work this paper cites.
Tree of Thoughts: Deliberate Problem Solving with Large Language Models
Shunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran, Thomas L. Griffiths, Yuan Cao, and Karthik Narasimhan. 2023 · 2023
Earlier work this paper cites.
Muxflow: Efficient and safe gpu sharing in large-scale production deep learning clusters
Yihao Zhao, Xin Liu, Shufan Liu, Xiang Li, Yibo Zhu, Gang Huang, Xuanzhe Liu, and Xin Jin. 2023 · 2023
Cited alongside, same era.
Response length perception and sequence scheduling: An llm-empowered llm inference pipeline
Zangwei Zheng, Xiaozhe Ren, Fuzhao Xue, Yang Luo, Xin Jiang, and Yang You. 2023b · 2023
Cited alongside, same era.
Rethinking the Uncertainty: A Critical Review and Analysis in the Era of Large Language Models
Mohammad Beigi, Sijia Wang, Ying Shen, Zihao Lin, Adithya Kulkarni, Jianfeng He, Feng Chen, Ming Jin, Jin-Hee Cho, Dawei Zhou, Chang-Tien Lu, and Lifu Huang. 2024 · 2024
Cited alongside, same era.
Prompt cache: Modular attention reuse for low-latency inference
In Gim, Guojun Chen, Seung-seob Lee, Nikhil Sarda, Anurag Khandelwal, and Lin Zhong. 2024 · 2024
Cited alongside, same era.
Load Balancers — Envoy Proxy Documentation
Envoy Project. 2025 · 2025
Closest in time.
GitHub Copilot: Your AI pair programmer
GitHub. 2025 · 2025
Closest in time.
About the Gateway API | GKE networking
Google Cloud. 2025a · 2025
Closest in time.
Compute Engine Reservations Overview
Google Cloud. 2025b · 2025
Closest in time.
Google Cloud Platform — Future-proof infrastructure. Powerful data and analytics. No ops, just code
Google Cloud Platform. 2025 · 2025
Closest in time.
Introducing Gemini: Your New Personal AI Assistant
Google LLC. 2025 · 2025
Closest in time.
Services, Load Balancing, and Networking
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Spotserve: Serving generative large language models on preemptible instances. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2 . 1112–1127
Xupeng Miao, Chunan Shi, Jiangfei Duan, Xiaoli Xi, Dahua Lin, Bin Cui, and Zhihao Jia. 2024 · 2024
Cited alongside, same era.
X-OpenMP—eXtreme fine-grained tasking using lock-less work stealing
Poornima Nookala, Kyle Chard, and Ioan Raicu. 2024 · 2024
Cited alongside, same era.
Flexllm: A system for co-serving large language model inference and parameter-efficient finetuning
Gabriele Oliaro, Xupeng Miao, Xinhao Cheng, Vineeth Kada, Ruohan Gao, Yingyi Huang, Remi Delacourt, April Yang, Yingcheng Wang, Mengdi Wu, et al · 2024
Cited alongside, same era.
Preble: Efficient Distributed Prompt Scheduling for LLM Serving
Vikranth Srivatsa, Zijian He, Reyna Abhyankar, Dongming Li, and Yiying Zhang. 2024 · 2024
Cited alongside, same era.
DynamoLLM: Designing LLM Inference Clusters for Performance and Energy Efficiency
Jovan Stojkovic, Chaojie Zhang, Íñigo Goiri, Josep Torrellas, and Esha Choukse. 2024 · 2024
Cited alongside, same era.
WildChat: 1M ChatGPT Interaction Logs in the Wild. In The Twelfth International Conference on Learning Representations
Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng. 2024 · 2024
Cited alongside, same era.
Sglang: Efficient execution of structured language model programs
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Livia Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E Gonzalez, et al · 2024
Cited alongside, same era.
SGLang v0.4: Zero-Overhead Batch Scheduler, Cache-Aware Load Balancer, Faster Structured Outputs
2024 · 2025
Cited alongside, same era.
Kubernetes Authors. 2025 · 2025
Closest in time.
SkyServe: Serving AI Models across Regions and Clouds with Spot Instances. In Proceedings of the Twentieth European Conference on Computer Systems (Rotterdam, Netherlands) (EuroSys ’25) . Association for Computing Machinery, New York, NY, USA, 159–175
Ziming Mao, Tian Xia, Zhanghao Wu, Wei-Lin Chiang, Tyler Griggs, Romil Bhardwaj, Zongheng Yang, Scott Shenker, and Ion Stoica. 2025 · 2025
Closest in time.
Behind the Scenes Scaling ChatGPT
Evan Morikawa. 2023 · 2025
Closest in time.
Hierarchical Autoscaling for Large Language Model Serving with Chiron
Archit Patke, Dhemath Reddy, Saurabh Jha, Chandra Narayanaswami, Zbigniew Kalbarczyk, and Ravishankar Iyer. 2025 · 2025
Closest in time.
Perplexity: AI-powered answer engine that provides accurate, trusted, and real-time answers to any question
Perplexity AI. 2025 · 2025
Closest in time.
Sam Altman: OpenAI Has Reached Roughly 800 Million Users
PYMNTS. 2025 · 2025
Closest in time.
Ray Serve: Scalable and Programmable Serving for ML Models
Ray Project. 2025 · 2025
Closest in time.
Intro to Ghostwriter
Replit Inc. 2025 · 2025
Closest in time.
Just four companies are hoarding tens of billions of dollars worth of Nvidia GPU chips
Sherwood News. 2024 · 2025
Closest in time.
TikTok owner ByteDance taps TSMC to make its own AI GPUs to stop relying on Nvidia
Anton Shilov. 2024 · 2025
Closest in time.
Wiretapping LLMs: Network Side-Channel Attacks on Interactive LLM Services
Mahdi Soleimani, Grace Jia, In Gim, Seung seob Lee, and Anurag Khandelwal. 2025 · 2025
Closest in time.
Amazon CodeWhisperer, Free for Individual Use, is Now Generally Available
Steve Roberts. 2023 · 2025
Closest in time.
Tabnine: AI Code Assistant
Tabnine Inc. 2025 · 2025
Closest in time.
Windsurf: The Most Powerful AI Code Editor
Windsurf. 2025 · 2025
Closest in time.
I Know What You Asked: Prompt Leakage via KV-Cache Sharing in Multi-Tenant LLM Serving. In NDSS
Guanlong Wu, Zheng Zhang, Yao Zhang, Weili Wang, Jianyu Niu, Ye Wu, and Yinqian Zhang. 2025 · 2025
Closest in time.
ServeGen: Workload Characterization and Generation of Large Language Model Serving in Production
Yuxing Xiang, Xue Li, Kun Qian, Wenyuan Yu, Ennan Zhai, and Xin Jin. 2025 · 2025
Closest in time.
Prism: Unleashing GPU Sharing for Cost-Efficient Multi-LLM Serving
Shan Yu, Jiarong Xing, Yifan Qiao, Mingyuan Ma, Yangmin Li, Yang Wang, Shuo Yang, Zhiqiang Xie, Shiyi Cao, Ke Bao, et al · 2025
Closest in time.