Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have resulted in a surging demand for planet-scale serving systems, where tens of thousands of GPUs continuously serve hundreds of millions of users.
Layer normalization, 2016
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton · 2016
Earlier work this paper cites.
Sigmoid-weighted linear units for neural network function approximation in reinforcement learning, 2017
Stefan Elfwing, Eiji Uchibe, and Kenji Doya · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
How many words do we read per minute? a review and meta-analysis of reading rate
Marc Brysbaert · 2019
Earlier work this paper cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism, 2019
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Mia Xu Chen, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Yonghui Wu, and Zhifeng Chen · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Rammer: Enabling holistic deep learning compiler optimizations with rTasks
Lingxiao Ma, Zhiqiang Xie, Zhi Yang, Jilong Xue, Youshan Miao, Wei Cui, Wenxiang Hu, Fan Yang, Lintao Zhang, and Lidong Zhou · 2020
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism, 2020
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2020
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism, 2020
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2020
Earlier work this paper cites.
Fastermoe: Modeling and optimizing training of large-scale dynamic pre-trained models
Jiaao He, Jidong Zhai, Tiago Antunes, Haojie Wang, Fuwen Luo, Shangfeng Shi, and Qin Li · 2022
Earlier work this paper cites.
Unity: Accelerating DNN training through joint optimization of algebraic transformations and parallelization
Colin Unger, Zhihao Jia, Wei Wu, Sina Lin, Mandeep Baines, Carlos Efrain Quintero Narvaez, Vinay Ramakrishnaiah, Nirmal Prajapati, Pat McCormick, Jamaludin Mohd-Yusof, Xi Luo, Dheevatsa Mudigere, Jongsoo Park, Misha Smelyanskiy, and Alex Aiken · 2022
Earlier work this paper cites.
Orca: A distributed serving system for { \{ Transformer-Based } \} generative models
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun · 2022
Earlier work this paper cites.
Alpa: Automating inter-and { \{ Intra-Operator } \} parallelism for distributed deep learning
Lianmin Zheng, Zhuohan Li, Hao Zhang, Yonghao Zhuang, Zhifeng Chen, Yanping Huang, Yida Wang, Yuanzhong Xu, Danyang Zhuo, Eric P Xing, et al · 2022
Earlier work this paper cites.
https://huggingface.co/datasets/anon8231489123/ShareGPT_Vicuna_unfiltered , 2023
Sharegpt · 2023
Earlier work this paper cites.
Sarathi: Efficient llm inference by piggybacking decodes with chunked prefills
Amey Agrawal, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S Gulavani, and Ramachandran Ramjee · 2023
Earlier work this paper cites.
GQA: Training generalized multi-query transformer models from multi-head checkpoints
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebron, and Sumit Sanghai · 2023
Earlier work this paper cites.
The desperate hunt for the a.i. boom’s most indispensable prize
Erin Griffith · 2023
Earlier work this paper cites.
Mistral 7b, 2023
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica · 2023
Earlier work this paper cites.
Alpaserve: Statistical multiplexing with model parallelism for deep learning serving
Zhuohan Li, Lianmin Zheng, Yinmin Zhong, Vincent Liu, Ying Sheng, Xin Jin, Yanping Huang, Zhifeng Chen, Hao Zhang, Joseph E Gonzalez, et al · 2023
Earlier work this paper cites.
Reinventing search with a new ai-powered microsoft bing and edge, your copilot for the web, May 2023
Yusuf Mehdi · 2023
Cited alongside, same era.
Introducing chatgpt, 2023
OpenAI · 2023
Cited alongside, same era.
Aspen: Breaking operator barriers for efficient parallelization of deep neural networks
Jongseok Park, Kyungmin Bin, Gibum Park, Sangtae Ha, and Kyunghan Lee · 2023
Cited alongside, same era.
The inference cost of search disruption – large language model cost analysis, February 2023
Dylan Patel and Afzal Ahmad · 2023
Cited alongside, same era.
Splitwise: Efficient generative llm inference using phase splitting, 2023
Pratyush Patel, Esha Choukse, Chaojie Zhang, Íñigo Goiri, Aashaka Shah, Saeed Maleki, and Ricardo Bianchini · 2023
Cited alongside, same era.
Efficiently scaling transformer inference
Reiner Pope, Sholto Douglas, Aakanksha Chowdhery, Jacob Devlin, James Bradbury, Jonathan Heek, Kefan Xiao, Shivani Agrawal, and Jeff Dean · 2023
Qserve: W4a8kv4 quantization and system co-design for efficient llm serving, 2024
Yujun Lin, Haotian Tang, Shang Yang, Zhekai Zhang, Guangxuan Xiao, Chuang Gan, and Song Han · 2024
Closest in time.
Build the future of ai with meta llama 3
Meta · 2024
Closest in time.
Deepspeed: Fastgen
Microsoft · 2024
Closest in time.
Nvidia dgx platform
NVIDIA · 2024
Closest in time.
Nvlink and nvlink switch
NVIDIA · 2024
Closest in time.
Tensorrt-llm
NVIDIA · 2024
Closest in time.
Tensorrt-llm
NVIDIA · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Flexgen: High-throughput generative inference of large language models with a single gpu
Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher Ré, Ion Stoica, and Ce Zhang · 2023
Cited alongside, same era.
Welder: Scheduling deep learning memory access via tile-graph
Yining Shi, Zhi Yang, Jilong Xue, Lingxiao Ma, Yuqing Xia, Ziming Miao, Yuxiao Guo, Fan Yang, and Lidong Zhou · 2023
Cited alongside, same era.
Introducing microsoft 365 copilot – your copilot for work, May 2023
Jared Spataro · 2023
Cited alongside, same era.
CUTLASS, January 2023
Vijay Thakkar, Pradeep Ramani, Cris Cecka, Aniket Shivam, Honghao Lu, Ethan Yan, Jack Kosaian, Mark Hoemmen, Haicheng Wu, Andrew Kerr, Matt Nicely, Duane Merrill, Dustyn Blasig, Fengqi Qiao, Piotr Majcher, Paul Springer, Markus Hohnerbach, Jin Wang, and Manish Gupta · 2023
Cited alongside, same era.
Tsmc sees ai chip output constraints lasting 1.5 years
CHENG TING-FANG · 2023
Cited alongside, same era.
Llama: Open and efficient foundation language models, 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Cited alongside, same era.
OpenAI · 2024
Closest in time.
Chatgpt’s weekly users have doubled in less than a year, 2024
Emma Roth · 2024
Closest in time.
ChatGPT Statistics (AUG 2024) - Users Growth Data — demandsage.com
Shubham Singh · 2024
Closest in time.
MLSys @ WukLab - Can Scheduling Overhead Dominate LLM Inference Performance? A Study of CPU Scheduling Overhead on Two Popular LLM Inference Systems — mlsys.wuklab.io
Vikranth Srivatsa, Dongming Li, Yiying Zhang, and Reyna Abhyankar · 2024
Closest in time.
Orion: Interference-aware, fine-grained gpu sharing for ml applications
Foteini Strati, Xianzhe Ma, and Ana Klimovic · 2024
Closest in time.
Quest: Query-aware sparsity for efficient long-context llm inference, 2024
Jiaming Tang, Yilong Zhao, Kan Zhu, Guangxuan Xiao, Baris Kasikci, and Song Han · 2024
Closest in time.
Forestcoll: Efficient collective communications on heterogeneous network fabrics, 2024
Liangyu Zhao, Saeed Maleki, Aashaka Shah, Ziyue Yang, Hossein Pourreza, and Arvind Krishnamurthy · 2024
Closest in time.
Sglang: Efficient execution of structured language model programs, 2024
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng · 2024
Closest in time.
Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving
Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang · 2024
Closest in time.
Intelligent router for llm workloads: Improving performance through workload-aware load balancing, 2025
Kunal Jain, Anjaly Parayil, Ankur Mallick, Esha Choukse, Xiaoting Qin, Jue Zhang, Íñigo Goiri, Rujia Wang, Chetan Bansal, Victor Rühle, Anoop Kulkarni, Steve Kofsky, and Saravan Rajmohan · 2025
Closest in time.
Hierarchical autoscaling for large language model serving with chiron, 2025
Archit Patke, Dhemath Reddy, Saurabh Jha, Chandra Narayanaswami, Zbigniew Kalbarczyk, and Ravishankar Iyer · 2025
Closest in time.
Aibrix: Towards scalable, cost-effective large language model inference infrastructure, 2025
The AIBrix Team, Jiaxin Shan, Varun Gupta, Le Xu, Haiyang Shi, Jingyuan Zhang, Ning Wang, Linhui Xu, Rong Kang, Tongping Liu, Yifei Zhang, Yiqing Zhu, Shuowei Jin, Gangmuk Lim, Binbin Chen, Zuzhi Chen, Xiao Liu, Xin Chen, Kante Yin, Chak-Pong Chung, Chenyu Jiang, Yicheng Lu, Jianjun Chen, Caixue Lin, Wu Xiang, Rui Shi, and Liguang Xie · 2025
Closest in time.