Fetching the paper…
Reading the bibliography…
DistServe improves the performance of large language models (LLMs) serving by disaggregating the prefill and decoding computation.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Ray: A distributed framework for emerging AI applications
Philipp Moritz, Robert Nishihara, Stephanie Wang, Alexey Tumanov, Richard Liaw, Eric Liang, Melih Elibol, Zongheng Yang, William Paul, Michael I. Jordan, and Ion Stoica · 2018
Earlier work this paper cites.
Fundamentals of queueing theory
John F Shortle, James M Thompson, Donald Gross, and Carl M Harris · 2018
Earlier work this paper cites.
Fastertransformer, 2019
NVIDIA Corporation · 2019
Earlier work this paper cites.
Triton inference server: An optimized cloud and edge inferencing solution., 2019
NVIDIA Corporation · 2019
Earlier work this paper cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism, 2019
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Mia Xu Chen, Dehao Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V. Le, Yonghui Wu, and Zhifeng Chen · 2019
Earlier work this paper cites.
Pipedream: Generalized pipeline parallelism for dnn training
Deepak Narayanan, Aaron Harlap, Amar Phanishayee, Vivek Seshadri, Nikhil R. Devanur, Gregory R. Ganger, Phillip B. Gibbons, and Matei Zaharia · 2019
Earlier work this paper cites.
Fast transformer decoding: One write-head is all you need, 2019
Noam Shazeer · 2019
Earlier work this paper cites.
Serving DNNs like clockwork: Performance predictability from the bottom up
Arpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao, Antoine Kaufmann, Ymir Vigfusson, and Jonathan Mace · 2020
Earlier work this paper cites.
Serving DNNs like clockwork: Performance predictability from the bottom up
Arpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao, Antoine Kaufmann, Ymir Vigfusson, and Jonathan Mace · 2020
Earlier work this paper cites.
Zero: Memory optimizations toward training trillion parameter models, 2020
Samyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, and Yuxiong He · 2020
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism, 2020
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2020
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
Turbotransformers: an efficient gpu serving system for transformer models
Jiarui Fang, Yang Yu, Chengduo Zhao, and Jie Zhou · 2021
Earlier work this paper cites.
Pollux: Co-adaptive cluster scheduling for goodput-optimized deep learning
Aurick Qiao, Sang Keun Choe, Suhas Jayaram Subramanya, Willie Neiswanger, Qirong Ho, Hao Zhang, Gregory R. Ganger, and Eric P. Xing · 2021
Earlier work this paper cites.
Ft-cnn: Algorithm-based fault tolerance for convolutional neural networks
Kai Zhao, Sheng Di, Sihuan Li, Xin Liang, Yujia Zhai, Jieyang Chen, Kaiming Ouyang, Franck Cappello, and Zizhong Chen · 2021
Earlier work this paper cites.
https://openai.com/blog/chatgpt , 2022
Introducing chatgpt · 2022
Earlier work this paper cites.
A case for disaggregation of ml data processing, 2022
Andrew Audibert, Yang Chen, Dan Graur, Ana Klimovic, Jiri Simsa, and Chandramohan A. Thekkath · 2022
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness, 2022
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré · 2022
Earlier work this paper cites.
Microsecond-scale preemption for concurrent GPU-accelerated DNN inferences
Mingcong Han, Hanze Zhang, Rong Chen, and Haibo Chen · 2022
Cited alongside, same era.
Orca: A distributed serving system for { \{ Transformer-Based } \} generative models
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun · 2022
Cited alongside, same era.
Opt: Open pre-trained transformer language models, 2022
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer · 2022
Cited alongside, same era.
Alpa: Automating inter- and Intra-Operator parallelism for distributed deep learning
Lianmin Zheng, Zhuohan Li, Hao Zhang, Yonghao Zhuang, Zhifeng Chen, Yanping Huang, Yida Wang, Yuanzhong Xu, Danyang Zhuo, Eric P. Xing, Joseph E. Gonzalez, and Ion Stoica · 2022
Cited alongside, same era.
{ \{ PetS } \} : A unified framework for { \{ Parameter-Efficient } \} transformers serving
Zhe Zhou, Xuechao Wei, Jiejing Zhang, and Guangyu Sun · 2022
Efficient memory management for large language model serving with pagedattention, 2023
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E. Gonzalez, Hao Zhang, and Ion Stoica · 2023
Later among the works it cites.
Alpaserve: Statistical multiplexing with model parallelism for deep learning serving
Zhuohan Li, Lianmin Zheng, Yinmin Zhong, Vincent Liu, Ying Sheng, Xin Jin, Yanping Huang, Zhifeng Chen, Hao Zhang, Joseph E Gonzalez, et al · 2023
Later among the works it cites.
Gpt-4 technical report, 2023
OpenAI · 2023
Later among the works it cites.
Splitwise: Efficient generative llm inference using phase splitting, 2023
Pratyush Patel, Esha Choukse, Chaojie Zhang, Íñigo Goiri, Aashaka Shah, Saeed Maleki, and Ricardo Bianchini · 2023
Later among the works it cites.
Reuters, 2023
2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
https://bard.google.com/ , 2023
Bard, an experiment by google · 2023
Cited alongside, same era.
Deepspeed model implementations for inference (mii), 2023
2023
Cited alongside, same era.
https://inflection.ai/assets/Inflection-1.pdf , 2023
Inflection tech memo · 2023
Cited alongside, same era.
Lanchain usecase: Summarization, 2023
2023
Cited alongside, same era.
Nvidia collective communications library (nccl), 2023
2023
Cited alongside, same era.
Serve, optimize and scale pytorch models in production, 2023
2023
Cited alongside, same era.
https://sharegpt.com/ , 2023
Sharegpt teams · 2023
Cited alongside, same era.
Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, et al · 2023
Later among the works it cites.
Flexgen: High-throughput generative inference of large language models with a single gpu, 2023
Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Daniel Y. Fu, Zhiqiang Xie, Beidi Chen, Clark Barrett, Joseph E. Gonzalez, Percy Liang, Christopher Ré, Ion Stoica, and Ce Zhang · 2023
Later among the works it cites.
Powerinfer: Fast large language model serving with a consumer-grade gpu, 2023
Yixin Song, Zeyu Mi, Haotong Xie, and Haibo Chen · 2023
Later among the works it cites.
Hotgpt: How to make software documentation more useful with a large language model?
Yiming Su, Chengcheng Wan, Utsav Sethi, Shan Lu, Madan Musuvathi, and Suman Nath · 2023
Later among the works it cites.
Llama: Open and efficient foundation language models, 2023
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample · 2023
Later among the works it cites.
Fast distributed inference serving for large language models
Bingyang Wu, Yinmin Zhong, Zili Zhang, Gang Huang, Xuanzhe Liu, and Xin Jin · 2023
Later among the works it cites.
Shepherd: Serving dnns in the wild
Hong Zhang, Yupeng Tang, Anurag Khandelwal, and Ion Stoica · 2023
Later among the works it cites.
Make it real: An end-to-end implementation of a physically disaggregated data center
Yiying Zhang · 2023
Later among the works it cites.
Introducing the next generation of claude
Anthropic · 2024
Closest in time.
Our next-generation model: Gemini 1.5
Google · 2024
Closest in time.
Inference without interference: Disaggregate llm inference for mixed downstream workloads, 2024
Cunchen Hu, Heyang Huang, Liangliang Xu, Xusheng Chen, Jiang Xu, Shuang Chen, Hao Feng, Chenxi Wang, Sa Wang, Yungang Bao, Ninghui Sun, and Yizhou Shan · 2024
Closest in time.
Distmind: Efficient resource disaggregation for deep learning workloads
Xin Jin, Zhihao Bai, Zhen Zhang, Yibo Zhu, Yinmin Zhong, and Xuanzhe Liu · 2024
Closest in time.
World model on million-length video and language with ringattention
Hao Liu, Wilson Yan, Matei Zaharia, and Pieter Abbeel · 2024
Closest in time.
Déjàvu: Kv-cache streaming for fast, fault-tolerant generative llm serving, 2024
Foteini Strati, Sara Mcallister, Amar Phanishayee, Jakub Tarnawski, and Ana Klimovic · 2024
Closest in time.