Fetching the paper…
Reading the bibliography…
This paper introduces Helix, a distributed system for high-throughput, low-latency large language model (LLM) serving in heterogeneous GPU clusters.
Analysis of preflow push algorithms for maximum network flow
Joseph Cheriyan and SN Maheshwari · 1989
Earlier work this paper cites.
Fast and effective task scheduling in heterogeneous systems
Andrei Radulescu and Arjan JC Van Gemund · 2000
Earlier work this paper cites.
Performance-effective and low-complexity task scheduling for heterogeneous computing
Haluk Topcuoglu, Salim Hariri, and Min-You Wu · 2002
Earlier work this paper cites.
Improving scheduling of tasks in a heterogeneous environment
Rashmi Bajaj and Dharma P Agrawal · 2004
Earlier work this paper cites.
Tensorflow-serving: Flexible, high-performance ml serving
Christopher Olston, Noah Fiedel, Kiril Gorovoy, Jeremiah Harmsen, Li Lao, Fangwei Li, Vinu Rajashekhar, Sukriti Ramesh, and Jordan Soyke · 2017
Earlier work this paper cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, et al · 2019
Earlier work this paper cites.
Parity models: erasure-coded resilience for prediction serving systems
Jack Kosaian, KV Rashmi, and Shivaram Venkataraman · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Fast transformer decoding: One write-head is all you need
Noam Shazeer · 2019
Earlier work this paper cites.
Nexus: A gpu cluster engine for accelerating dnn-based video analysis
Haichen Shen, Lequn Chen, Yuchen Jin, Liangyu Zhao, Bingyu Kong, Matthai Philipose, Arvind Krishnamurthy, and Ravi Sundaram · 2019
Earlier work this paper cites.
Megatron-lm: Training multi-billion parameter language models using model parallelism
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Serving { \{ DNNs } \} like clockwork: Performance predictability from the bottom up
Arpan Gujarati, Reza Karimi, Safya Alzayat, Wei Hao, Antoine Kaufmann, Ymir Vigfusson, and Jonathan Mace · 2020
Earlier work this paper cites.
Learning-based coded computation
Jack Kosaian, KV Rashmi, and Shivaram Venkataraman · 2020
Earlier work this paper cites.
Collage inference: Using coded redundancy for lowering latency variation in distributed image classification systems
Hema Venkata Krishna Giri Narra, Zhifeng Lin, Ganesh Ananthanarayanan, Salman Avestimehr, and Murali Annavaram · 2020
Earlier work this paper cites.
Towards crowdsourced training of large neural networks using decentralized mixture-of-experts
Max Ryabinin and Anton Gusev · 2020
Earlier work this paper cites.
Distributed deep learning in open collaborations
Michael Diskin, Alexey Bukhtiyarov, Max Ryabinin, Lucile Saulnier, Anton Sinitsin, Dmitry Popov, Dmitry V Pyrkin, Maxim Kashirin, Alexander Borzunov, Albert Villanova del Moral, et al · 2021
Earlier work this paper cites.
Interleaved weighted round-robin: A network calculus analysis
Seyed Mohammadhossein Tabatabaee, Jean-Yves Le Boudec, and Marc Boyer · 2021
Earlier work this paper cites.
Petals: Collaborative inference and fine-tuning of large models
Alexander Borzunov, Dmitry Baranchuk, Tim Dettmers, Max Ryabinin, Younes Belkada, Artem Chumachenko, Pavel Samygin, and Colin Raffel · 2022
Earlier work this paper cites.
Whale: Efficient giant model training over heterogeneous { \{ GPUs } \}
Xianyan Jia, Le Jiang, Ang Wang, Wencong Xiao, Ziji Shi, Jie Zhang, Xinyuan Li, Langshi Chen, Yong Li, Zhen Zheng, et al · 2022
Earlier work this paper cites.
Orca: A distributed serving system for { \{ Transformer-Based } \} generative models
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun · 2022
Cited alongside, same era.
Decentralized training of foundation models in heterogeneous environments
Binhang Yuan, Yongjun He, Jared Davis, Tianyi Zhang, Tri Dao, Beidi Chen, Percy S Liang, Christopher Re, and Ce Zhang · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Sarathi: Efficient llm inference by piggybacking decodes with chunked prefills
Amey Agrawal, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S Gulavani, and Ramachandran Ramjee · 2023
Cited alongside, same era.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints
Transparent { \{ GPU } \} sharing in container clouds for deep learning workloads
Bingyang Wu, Zili Zhang, Zhihao Bai, Xuanzhe Liu, and Xin Jin · 2023
Later among the works it cites.
{ \{ SkyPilot } \} : An intercloud broker for sky computing
Zongheng Yang, Zhanghao Wu, Michael Luo, Wei-Lin Chiang, Romil Bhardwaj, Woosuk Kwon, Siyuan Zhuang, Frank Sifei Luan, Gautam Mittal, Scott Shenker, et al · 2023
Later among the works it cites.
Distributed inference and fine-tuning of large language models over the internet
Alexander Borzunov, Max Ryabinin, Artem Chumachenko, Dmitry Baranchuk, Tim Dettmers, Younes Belkada, Pavel Samygin, and Colin A Raffel · 2024
Closest in time.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Closest in time.
Google cloud compute products
Google Cloud · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai · 2023
Cited alongside, same era.
Gurobi Optimizer Reference Manual, 2023
Gurobi Optimization, LLC · 2023
Cited alongside, same era.
Metagpt: Meta programming for multi-agent collaborative framework
Sirui Hong, Xiawu Zheng, Jonathan Chen, Yuheng Cheng, Jinlin Wang, Ceyao Zhang, Zili Wang, Steven Ka Shing Yau, Zijuan Lin, Liyang Zhou, et al · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica · 2023
Cited alongside, same era.
Fast inference from transformers via speculative decoding
Yaniv Leviathan, Matan Kalman, and Yossi Matias · 2023
Cited alongside, same era.
{ \{ AlpaServe } \} : Statistical multiplexing with model parallelism for deep learning serving
Zhuohan Li, Lianmin Zheng, Yinmin Zhong, Vincent Liu, Ying Sheng, Xin Jin, Yanping Huang, Zhifeng Chen, Hao Zhang, Joseph E Gonzalez, et al · 2023
Cited alongside, same era.
Wizardcoder: Empowering code large language models with evol-instruct
Ziyang Luo, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang · 2023
Cited alongside, same era.
Towards efficient generative large language model serving: A survey from algorithms to systems
Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Hongyi Jin, Tianqi Chen, and Zhihao Jia · 2023
Cited alongside, same era.
Tyler Griggs, Xiaoxuan Liu, Jiaxiang Yu, Doyoung Kim, Wei-Lin Chiang, Alvin Cheung, and Ion Stoica · 2024
Closest in time.
Hexgen: Generative inference of large language model over heterogeneous environment
Youhe Jiang, Ran Yan, Xiaozhe Yao, Yang Zhou, Beidi Chen, and Binhang Yuan · 2024
Closest in time.
Energy aware scheduling
Linux Kernel Documentation · 2024
Closest in time.
Realhf: Optimized rlhf training for large language models through parameter reallocation, 2024
Zhiyu Mei, Wei Fu, Kaiwei Li, Guangju Wang, Huanchen Zhang, and Yi Wu · 2024
Closest in time.
Introducing Meta Llama 3: The most capable openly available LLM to date — ai.meta.com
Meta · 2024
Closest in time.
Specinfer: Accelerating large language model serving with tree-based speculative inference and verification
Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Zeyu Wang, Zhengxin Zhang, Rae Ying Yee Wong, Alan Zhu, Lijie Yang, Xiaoxiang Shi, et al · 2024
Closest in time.
NVIDIA A100 Tensor Core GPU
NVIDIA · 2024
Closest in time.
NVIDIA H100 Tensor Core GPU
NVIDIA · 2024
Closest in time.
NVIDIA L4 Tensor Core GPU
NVIDIA · 2024
Closest in time.
NVIDIA T4 Tensor Core GPU
NVIDIA · 2024
Closest in time.
Splitwise: Efficient generative llm inference using phase splitting
Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, and Ricardo Bianchini · 2024
Closest in time.
Ml training with cloud gpu shortages: Is cross-region the answer?
Foteini Strati, Paul Elvinger, Tolga Kerimoglu, and Ana Klimovic · 2024
Closest in time.
ZeroMQ An open-source universal messaging library, 2024
The ZeroMQ authors · 2024
Closest in time.
Hap: Spmd dnn training on heterogeneous gpu clusters with automated program synthesis
Shiwei Zhang, Lansong Diao, Chuan Wu, Zongyan Cao, Siyu Wang, and Wei Lin · 2024
Closest in time.
{ \{ DistServe } \} : Disaggregating prefill and decoding for goodput-optimized large language model serving
Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang · 2024
Closest in time.