Fetching the paper…
Reading the bibliography…
Serving disaggregated large language models (LLMs) over tens of thousands of xPU devices (GPUs or NPUs) with reliable performance faces multiple challenges.
https://github.com/apache/zookeeper , 2010
Apache Zookeeper · 2010
Earlier work this paper cites.
https://github.com/kubernetes/kubernetes , 2014
Kubernetes · 2014
Earlier work this paper cites.
Attention is all you need, 2017
Vaswani Ashish, Shazeer Noam, Parmar Niki, Uszkoreit Jakob, Jones Llion, N. Gomez Aidan, Kaiser Lukasz, and Polosukhin Illia · 2017
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova · 2018
Earlier work this paper cites.
https://github.com/Ascend/ascend-for-volcano , 2019
Ascend-for-volcano · 2019
Earlier work this paper cites.
https://github.com/NVIDIA/FasterTransformer , 2021
Nvidia, FasterTransformer · 2021
Earlier work this paper cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei Jason, Wang Xuezhi, Schuurmans Dale, Bosma Maarten, Ichter Brian, Xia Fei, H. Chi Ed, V. Le Quoc, and Zhou Denny · 2022
Earlier work this paper cites.
Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale
Reza Yazdani Aminabadi, Samyam Rajbhandari, Ammar Ahmad Awan, Cheng Li, Du Li, Elton Zheng, Olatunji Ruwase, Shaden Smith, Minjia Zhang, Jeff Rasley, et al · 2022
Earlier work this paper cites.
Orca: A distributed serving system for transformer-based generative models
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun · 2022
Earlier work this paper cites.
Gpt3. int8 (): 8-bit matrix multiplication for transformers at scale
Tim Dettmers, Mike Lewis, Younes Belkada, and Luke Zettlemoyer · 2022
Earlier work this paper cites.
https://huggingface.co/blog/bloom-megatron-deepspeed , 2022
The Technology Behind BLOOM Training · 2022
Earlier work this paper cites.
Gemini: a family of highly capable multimodal models
Gemini Team, Rohan Anil, Sebastian Borgeaud, Yonghui Wu, Jean-Baptiste Alayrac, Jiahui Yu, Radu Soricut, Johan Schalkwyk, Andrew M Dai, Anja Hauth, et al · 2023
Earlier work this paper cites.
Pangu- Σ \Sigma : Towards trillion parameter language model with sparse heterogeneous computing
Ren Xiaozhe, Zhou Pingyi, Meng Xinfan, Huang Xinjing, Wang Yadao, Wang Weichao, Li Pengfei, Zhang Xiaoda, Podolskiy Alexander, Arshinov Grigory, Bout Andrey, Piontkovskaya Irina, Wei Jiansheng, Jiang Xin, Su Teng, Liu Qun, and Yao Jun · 2023
Earlier work this paper cites.
https://kimi.moonshot.cn , 2023
Moonshot AI · 2023
Earlier work this paper cites.
Splitwise: Efficient generative llm inference using phase splitting
Pratyush Patel, Esha Choukse, Chaojie Zhang, Íñigo Goiri, Aashaka Shah, Saeed Maleki, and Ricardo Bianchini · 2023
Earlier work this paper cites.
vllm: Easy, fast, and cheap llm serving with pagedattention
vLLM Project · 2023
Earlier work this paper cites.
Prompt cache: Modular attention reuse for low-latency inference
In Gim, Guojun Chen, Seung-seob Lee, Nikhil Sarda, Anurag Khandelwal, and Lin Zhong · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Kwon Woosuk, Li Zhuohan, Zhuang Siyuan, Sheng Ying, Zheng Lianmin, Hao Yu Cody, E. Gonzalez Joseph, Zhang Hao, and Stoica Ion · 2023
Earlier work this paper cites.
A prompt pattern catalog to enhance prompt engineering with chatgpt
White Jules, Fu Quchen, Hays Sam, Sandborn Michael, Olea Carlos, Gilbert Henry, Elnashar Ashraf, Spencer-Smith Jesse, and C. Schmidt Douglas · 2023
Cited alongside, same era.
Sarathi: Efficient llm inference by piggybacking decodes with chunked prefills
Amey Agrawal, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S Gulavani, and Ramachandran Ramjee · 2023
Cited alongside, same era.
Flashattention-2: Faster attention with better parallelism and work partitioning
Tri Dao · 2023
Cited alongside, same era.
Fast distributed inference serving for large language models
Bingyang Wu, Yinmin Zhong, Zili Zhang, Gang Huang, Xuanzhe Liu, and Xin Jin · 2023
Cited alongside, same era.
https://openai.com/index/gpt-4 , 2024
OpenAI · 2024
Closest in time.
Inference without interference: Disaggregate llm inference for mixed downstream workloads
Hu Cunchen, Huang Heyang, Xu Liangliang, Chen Xusheng, Xu Jiang, Chen Shuang, Feng Hao, Wang Chenxi, Wang Sa, Bao Yungang, Sun Ninghui, and Shan Yizhou · 2024
Closest in time.
Cost-efficient large language model serving for multi-turn conversations with cachedattention
Gao Bin, He Zhuomin, Sharma Puru, Kang Qingxuan, Jevdjic Djordje, Deng Junbo, Yang Xingkun, Yu Zhou, and Zuo Pengfei · 2024
Closest in time.
Mooncake: A kvcache-centric disaggregated architecture for llm serving
Qin Ruoyu, Li Zheming, He Weiran, Zhang Mingxing, Wu Yongwei, Zheng Weimin, and Xu Xinran · 2024
Closest in time.
Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Li Zhuohan, Zheng Lianmin, Zhong Yinmin, Liu Vincent, Sheng Ying, Jin Xin, Huang Yanping, Chen Zhifeng, Zhang Hao, E. Gonzalez Joseph, and Stoica Ion · 2023
Cited alongside, same era.
Spqr: A sparse-quantized representation for near-lossless llm weight compression
Tim Dettmers, Ruslan Svirschevski, Vage Egiazarian, Denis Kuznedelev, Elias Frantar, Saleh Ashkboos, Alexander Borzunov, Torsten Hoefler, and Dan Alistarh · 2023
Cited alongside, same era.
Gpt-zip: Deep compression of finetuned large language models
Berivan Isik, Hermann Kumbong, Wanyi Ning, Xiaozhe Yao, Sanmi Koyejo, and Ce Zhang · 2023
Cited alongside, same era.
Smoothquant: Accurate and efficient post-training quantization for large language models
Guangxuan Xiao, Ji Lin, Mickael Seznec, Hao Wu, Julien Demouth, and Song Han · 2023
Cited alongside, same era.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai · 2023
Cited alongside, same era.
Llamarec: Two-stage recommendation using large language models for ranking
Yue Zhenrui, Rabhi Sara, Moreira Gabriel, de Souza Pereira, Wang Dong, and Oldridge Even · 2023
Cited alongside, same era.
https://github.com/Ascend/ascend-device-plugin , 2023
Ascend Device Plugin · 2023
Cited alongside, same era.
https://github.com/NVIDIA/TensorRT-LLM , 2023
Nvidia, TensorRT-LLM · 2023
Cited alongside, same era.
Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang · 2024
Closest in time.
Chunkattention: Efficient self-attention with prefix-aware kv cache and two-phase partition
Lu Ye, Ze Tao, Yong Huang, and Yang Li · 2024
Closest in time.
Taming throughput-latency tradeoff in llm inference with sarathi-serve
Agrawal Amey, Kedia Nitin, Panwar Ashish, Mohan Jayashree, Kwatra Nipun, S. Gulavani Bhargav, Tumanov Alexey, and Ramjee Ramachandran · 2024
Closest in time.
Fairness in serving large language models
Sheng Ying, Cao Shiyi, Li Dacheng, Zhu Banghua, Li Zhuohan, Zhuo Danyang, E. Gonzalez Joseph, and Stoica Ion · 2024
Closest in time.
A survey on transformer compression
Tang Yehui, Wang Yunhe, Guo Jianyuan, Tu Zhijun, Han Kai, Hu Hailin, and Tao Dacheng · 2024
Closest in time.
Loongserve: Efficiently serving long-context large language models with elastic sequence parallelism
Wu Bingyang, Liu Shengyu, Zhong Yinmin, Sun Peng, Liu Xuanzhe, and Jin Xin · 2024
Closest in time.
A survey on rag meeting llms: Towards retrieval-augmented large language models
Fan Wenqi, Ding Yujuan, Ning Liangbo, Wang Shijie, Li Hengyun, Yin Dawei, Chua Tat-Seng, and Li Qing · 2024
Closest in time.
High-speed Interfaces (HCCS, PCIe, RoCE) in the Huawei Cluster
Atlas AI Cluster · 2024
Closest in time.
Déjàvu: Kv-cache streaming for fast, fault-tolerant generative llm serving
Foteini Strati, Sara Mcallister, Amar Phanishayee, Jakub Tarnawski, and Ana Klimovic · 2024
Closest in time.
Memserve: Context caching for disaggregated llm serving with elastic memory pool
Hu Cunchen, Huang Heyang, Hu Junhao, Xu Jiang, Chen Xusheng, Xie Tao, Wang Chenxi, Wang Sa, Bao Yungang, Sun Ninghui, and Shan Yizhou · 2024
Closest in time.
Specinfer: Accelerating large language model serving with tree-based speculative inference and verification
Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Zeyu Wang, Zhengxin Zhang, Ying Yee Wong Rae, Alan Zhu, Lijie Yang, Xiaoxiang Shi, Chunan Shi, Zhuoming Chen, Daiyaan Arfeen, Reyna Abhyankar, and Zhihao Jia · 2024
Closest in time.
Flashdecoding++: Faster large language model inference on gpus
Ke Hong, Guohao Dai, Jiaming Xu, Qiuli Mao, Xiuhong Li, Jun Liu, Kangdi Chen, Yuhan Dong, and Yu Wang · 2024
Closest in time.
Muxserve: Flexible spatial-temporal multiplexing for multiple llm serving
Jiangfei Duan, Runyu Lu, Haojie Duanmu, Xiuhong Li, Xingcheng Zhang, Dahua Lin, Ion Stoica, and Hao Zhang · 2024
Closest in time.