Fetching the paper…
Reading the bibliography…
Inference on large-language models (LLMs) is constrained by GPU memory capacity.
Salus: Fine-Grained GPU Sharing Primitives for Deep Learning Applications
Peifeng Yu and Mosharaf Chowdhury. 2019 · 1902
Earlier work this paper cites.
Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2020 · 1909
Earlier work this paper cites.
A Neural Probabilistic Language Model. In Advances in Neural Information Processing Systems , T. Leen, T. Dietterich, and V. Tresp (Eds.), Vol. 13. MIT Press
Yoshua Bengio, Réjean Ducharme, and Pascal Vincent. 2000 · 2000
Earlier work this paper cites.
Efficient Estimation of Word Representations in Vector Space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 · 2013
Earlier work this paper cites.
Reducing Internet Latency: A Survey of Techniques and Their Merits
Bob Briscoe, Anna Brunstrom, Andreas Petlund, David Hayes, David Ros, Ing-Jyh Tsang, Stein Gjessing, Gorry Fairhurst, Carsten Griwodz, and Michael Welzl. 2016 · 2014
Earlier work this paper cites.
Jupiter Rising: A Decade of Clos Topologies and Centralized Control in Google’s Datacenter Network. In Sigcomm ’15
Arjun Singh, Joon Ong, Amit Agarwal, Glen Anderson, Ashby Armistead, Roy Bannon, Seb Boving, Gaurav Desai, Bob Felderman, Paulie Germano, Anand Kanagala, Jeff Provost, Jason Simmons, Eiichi Tanda, Jim Wanderer, Urs Hölzle, Stephen Stuart, and Amin Vahdat. 2015 · 2015
Earlier work this paper cites.
Efficient Memory Disaggregation with Infiniswap. In 14th USENIX Symposium on Networked Systems Design and Implementation (NSDI 17) . USENIX Association, Boston, MA, 649–667
Juncheng Gu, Youngmoon Lee, Yiwen Zhang, Mosharaf Chowdhury, and Kang G. Shin. 2017 · 2017
Earlier work this paper cites.
Attention is All you Need. In Advances in Neural Information Processing Systems , I. Guyon, U. Von Luxburg, S. Bengio, H. Wallach, R. Fergus, S. Vishwanathan, and R. Garnett (Eds.), Vol. 30. Curran Associates, Inc
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Gandiva: Introspective Cluster Scheduling for Deep Learning. In 13th USENIX Symposium on Operating Systems Design and Implementation (OSDI 18) . USENIX Association, Carlsbad, CA, 595–610
Wencong Xiao, Romil Bhardwaj, Ramachandran Ramjee, Muthian Sivathanu, Nipun Kwatra, Zhenhua Han, Pratyush Patel, Xuan Peng, Hanyu Zhao, Quanlu Zhang, Fan Yang, and Lidong Zhou. 2018 · 2018
Earlier work this paper cites.
BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Can far memory improve job throughput?. In Proceedings of the Fifteenth European Conference on Computer Systems (Heraklion, Greece) (EuroSys ’20) . Association for Computing Machinery, New York, NY, USA, Article 14, 16 pages
Emmanuel Amaro, Christopher Branner-Augmon, Zhihong Luo, Amy Ousterhout, Marcos K. Aguilera, Aurojit Panda, Sylvia Ratnasamy, and Scott Shenker. 2020 · 2020
Earlier work this paper cites.
PipeSwitch: Fast Pipelined Context Switching for Deep Learning Applications. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20) . USENIX Association, 499–514
Zhihao Bai, Zhen Zhang, Yibo Zhu, and Xin Jin. 2020 · 2020
Earlier work this paper cites.
Serving DNNs like Clockwork: Performance Predictability from the Bottom Up. In 14th USENIX Symposium on Operating Systems Design and Implementation (OSDI 20) . USENIX Association, 443–462
Arpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao, Antoine Kaufmann, Ymir Vigfusson, and Jonathan Mace. 2020 · 2020
Earlier work this paper cites.
The future of AI is Wafer-Scale
Cerebras 2021 · 2021
Earlier work this paper cites.
Intel Gaudi AI accelerator
Intel Gaudi AI accelerator 2021 · 2021
Earlier work this paper cites.
Nvidia DGX Systems
Nvidia DGX Systems 2021 · 2021
Earlier work this paper cites.
DeepSpeed-inference: enabling efficient inference of transformer models at unprecedented scale. In Proceedings of the International Conference on High Performance Computing, Networking, Storage and Analysis (Dallas, Texas) (SC ’22) . IEEE Press, Article 46, 15 pages
Reza Yazdani Aminabadi, Samyam Rajbhandari, Ammar Ahmad Awan, Cheng Li, Du Li, Elton Zheng, Olatunji Ruwase, Shaden Smith, Minjia Zhang, Jeff Rasley, and Yuxiong He. 2022 · 2022
Earlier work this paper cites.
Memory Harvesting in Multi-GPU Systems with Hierarchical Unified Virtual Memory. In 2022 USENIX Annual Technical Conference (USENIX ATC 22) . USENIX Association, Carlsbad, CA, 625–638
Sangjin Choi, Taeksoo Kim, Jinwoo Jeong, Rachata Ausavarungnirun, Myeongjae Jeon, Youngjin Kwon, and Jeongseob Ahn. 2022 · 2022
Earlier work this paper cites.
High-Resolution Image Synthesis with Latent Diffusion Models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022 · 2022
Earlier work this paper cites.
Orca: A Distributed Serving System for Transformer-Based Generative Models. In 16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22) . USENIX Association, Carlsbad, CA, 521–538
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun. 2022 · 2022
Earlier work this paper cites.
Host Congestion Control. In Proceedings of the ACM SIGCOMM 2023 Conference (<conf-loc>, <city>New York</city>, <state>NY</state>, <country>USA</country>, </conf-loc>) (ACM SIGCOMM ’23) . Association for Computing Machinery, New York, NY, USA, 275–287
Saksham Agarwal, Arvind Krishnamurthy, and Rachit Agarwal. 2023 · 2023
Cited alongside, same era.
Norman P. Jouppi, George Kurian, Sheng Li, Peter Ma, Rahul Nagarajan, Lifeng Nai, Nishant Patil, Suvinay Subramanian, Andy Swing, Brian Towles, Cliff Young, Xiang Zhou, Zongwei Zhou, and David Patterson. 2023 · 2023
Cited alongside, same era.
AudioGen: Textually Guided Audio Generation
Felix Kreuk, Gabriel Synnaeve, Adam Polyak, Uriel Singer, Alexandre Défossez, Jade Copet, Devi Parikh, Yaniv Taigman, and Yossi Adi. 2023 · 2023
Cited alongside, same era.
NVLink & NVSwitch: Fastest HPC Data Center Platform
NVIDIA. 2024 · 2024
Closest in time.
NVIDIA B200
NVIDIA Corporation. 2024a · 2024
Closest in time.
NVIDIA DGX A100 Datasheet
NVIDIA Corporation. 2024b · 2024
Closest in time.
NVIDIA H100 Tensor Core GPU Datasheet
NVIDIA Corporation. 2024c · 2024
Closest in time.
OpenAI API
OpenAI. 2024a · 2024
Closest in time.
OpenAI batch API
OpenAI. 2024b · 2024
Closest in time.
OpenAI Customer Stories
OpenAI. 2024c · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Efficient Memory Management for Large Language Model Serving with PagedAttention. In Proceedings of the 29th Symposium on Operating Systems Principles (<conf-loc>, <city>Koblenz</city>, <country>Germany</country>, </conf-loc>) (SOSP ’23) . Association for Computing Machinery, New York, NY, USA, 611–626
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023 · 2023
Cited alongside, same era.
FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU
Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Daniel Y. Fu, Zhiqiang Xie, Beidi Chen, Clark Barrett, Joseph E. Gonzalez, Percy Liang, Christopher Ré, Ion Stoica, and Ce Zhang. 2023 · 2023
Cited alongside, same era.
PyTorch: An open source machine learning framework that accelerates the path from research prototyping to production deployment
2024 · 2024
Cited alongside, same era.
Together: The fastest cloud platform for building and running generative AI
Together AI. 2024c · 2024
Cited alongside, same era.
EC2 Auto Scaling Warm Pools
Amazon Web Services. 2024 · 2024
Cited alongside, same era.
Prompt Caching
Anthropic. [n. d.] · 2024
Cited alongside, same era.
Arxiv summarization dataset
arXiv. 2023 · 2024
Cited alongside, same era.
Generative AI-Use Cases and Resources
Amazon Web Services (AWS). 2024 · 2024
Cited alongside, same era.
Azure AI Studio
Microsoft Azure. [n. d.] · 2024
Cited alongside, same era.
Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, and Ricardo Bianchini. 2024 · 2024
Closest in time.
Amazon SageMaker
Amazon Web Services. [n. d.] · 2024
Closest in time.
Mewtant Case Study
Amazon Web Services. 2024 · 2024
Closest in time.
Fairness in Serving Large Language Models. In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) . USENIX Association, Santa Clara, CA, 965–988
Ying Sheng, Shiyi Cao, Dacheng Li, Banghua Zhu, Zhuohan Li, Danyang Zhuo, Joseph E. Gonzalez, and Ion Stoica. 2024 · 2024
Closest in time.
Llumnix: Dynamic Scheduling for Large Language Model Serving. In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) . USENIX Association, Santa Clara, CA, 173–191
Biao Sun, Ziming Huang, Hanyu Zhao, Wencong Xiao, Xinyi Zhang, Yong Li, and Wei Lin. 2024 · 2024
Closest in time.
CFS Scheduler
The Linux Kernel Developers. [n. d.] · 2024
Closest in time.
A case for server-scale photonic connectivity. In Proceedings of the 23rd ACM Workshop on Hot Topics in Networks (Irvine, CA, USA) (HotNets ’24) . Association for Computing Machinery, New York, NY, USA, 290–299
Abhishek Vijaya Kumar, Arjun Devraj, Darius Bunandar, and Rachee Singh. 2024 · 2024
Closest in time.
Distributed LLM serving
vLLM. 2023 · 2024
Closest in time.
LLama31: Enhancing LLM Serving Efficiency
VLLM. 2024 · 2024
Closest in time.
Fast Distributed Inference Serving for Large Language Models
Bingyang Wu, Yinmin Zhong, Zili Zhang, Shengyu Liu, Fangyue Liu, Yuanhang Sun, Gang Huang, Xuanzhe Liu, and Xin Jin. 2024 · 2024
Closest in time.
HeteGen: Heterogeneous Parallel Inference for Large Language Models on Resource-Constrained Devices
Xuanlei Zhao, Bin Jia, Haotian Zhou, Ziming Liu, Shenggan Cheng, and Yang You. 2024 · 2024
Closest in time.
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving. In 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24) . USENIX Association, Santa Clara, CA, 193–210
Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang. 2024 · 2024
Closest in time.
Azure LLM Inference Dataset 2024
Microsoft Azure. 2024 · 2025
Closest in time.