Fetching the paper…
Reading the bibliography…
The context window of large language models (LLMs) is rapidly increasing, leading to a huge variance in resource usage between different requests as well as between different phases of the same request.
Efficient dynamic programming using quadrangle inequalities. In ACM Symposium on Theory of Computing
F. Frances Yao. 1980 · 1980
Earlier work this paper cites.
Attention is all you need. In Neural Information Processing Systems
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Ray: A Distributed Framework for Emerging AI Applications. In USENIX OSDI
Philipp Moritz, Robert Nishihara, Stephanie Wang, Alexey Tumanov, Richard Liaw, Eric Liang, Melih Elibol, Zongheng Yang, William Paul, Michael I. Jordan, and Ion Stoica. 2018 · 2018
Earlier work this paper cites.
Optimus: an efficient dynamic resource scheduler for deep learning clusters. In EuroSys
Yanghua Peng, Yixin Bao, Yangrui Chen, Chuan Wu, and Chuanxiong Guo. 2018 · 2018
Earlier work this paper cites.
Generating Long Sequences with Sparse Transformers. In arXiv
Rewon Child, Scott Gray, Alec Radford, and Ilya Sutskever. 2019 · 2019
Earlier work this paper cites.
Tiresias: A GPU Cluster Manager for Distributed Deep Learning.. In USENIX NSDI
Juncheng Gu, Mosharaf Chowdhury, Kang G Shin, Yibo Zhu, Myeongjae Jeon, Junjie Qian, Hongqiang Harry Liu, and Chuanxiong Guo. 2019 · 2019
Earlier work this paper cites.
Fast Transformer Decoding: One Write-Head is All You Need. In arXiv
Noam Shazeer. 2019 · 2019
Earlier work this paper cites.
Megatron-LM: Training Multi-billion Parameter Language Models using Model Parallelism. In arXiv
Mohammad Shoeybi, Mostofa Patwary, Raul Puri, Patrick LeGresley, Jared Casper, and Bryan Catanzaro. 2019 · 2019
Earlier work this paper cites.
Triton: an intermediate language and compiler for tiled neural network computations. In Proceedings of the 3rd ACM SIGPLAN International Workshop on Machine Learning and Programming Languages
Philippe Tillet, H. T. Kung, and David Cox. 2019 · 2019
Earlier work this paper cites.
Elastic Resource Sharing for Distributed Deep Learning. In USENIX NSDI
Changho Hwang, Taehyun Kim, Sunghyun Kim, Jinwoo Shin, and KyoungSoo Park. 2021 · 2021
Earlier work this paper cites.
Adaptive Elastic Training for Sparse Deep Learning on Heterogeneous Multi-GPU Servers. In arXiv
Yujing Ma, Florin Rusu, Kesheng Wu, and Alexander Sim. 2021 · 2021
Earlier work this paper cites.
Pollux: Co-adaptive Cluster Scheduling for Goodput-Optimized Deep Learning. In USENIX OSDI
Aurick Qiao, Sang Keun Choe, Suhas Jayaram Subramanya, Willie Neiswanger, Qirong Ho, Hao Zhang, Gregory R. Ganger, and Eric P. Xing. 2021 · 2021
Earlier work this paper cites.
Introducing ChatGPT
2022 · 2022
Earlier work this paper cites.
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness. In Neural Information Processing Systems
Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. 2022 · 2022
Earlier work this paper cites.
Reducing Activation Recomputation in Large Transformer Models. In arXiv
Vijay Korthikanti, Jared Casper, Sangkug Lym, Lawrence McAfee, Michael Andersch, Mohammad Shoeybi, and Bryan Catanzaro. 2022 · 2022
Earlier work this paper cites.
Orca: A Distributed Serving System for Transformer-BasedGenerative Models. In USENIX OSDI
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun. 2022 · 2022
Earlier work this paper cites.
Bard, an experiment by Google
2023 · 2023
Earlier work this paper cites.
Optimized primitives for collective multi-GPU communicatio Resources
2023 · 2023
Earlier work this paper cites.
ShareGPT Teams
2023 · 2023
Cited alongside, same era.
GPT-4 technical report. In arXiv
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills. In arXiv
Amey Agrawal, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S. Gulavani, and Ramachandran Ramjee. 2023 · 2023
Cited alongside, same era.
GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints. In The Conference on Empirical Methods in Natural Language Processing (EMNLP)
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebron, and Sumit Sanghai. 2023 · 2023
Cited alongside, same era.
L-Eval: Instituting Standardized Evaluation for Long Context Language Models. In arXiv
Chenxin An, Shansan Gong, Ming Zhong, Xingjian Zhao, Mukai Li, Jun Zhang, Lingpeng Kong, and Xipeng Qiu. 2023 · 2023
Cited alongside, same era.
Fast Distributed Inference Serving for Large Language Models. In arXiv
Bingyang Wu, Yinmin Zhong, Zili Zhang, Gang Huang, Xuanzhe Liu, and Xin Jin. 2023 · 2023
Later among the works it cites.
DeepSpeed Model Implementations for Inference (MII)
2024 · 2024
Closest in time.
Keyformer: KV Cache Reduction through Key Tokens Selection for Efficient Generative Inference. In Conference on Machine Learning and Systems
Muhammad Adnan, Akhil Arunkumar, Gaurav Jain, Prashant J Nair, Ilya Soloveychik, and Purushotham Kamath. 2024 · 2024
Closest in time.
Introducing the next generation of Claude
Anthropic. 2024 · 2024
Closest in time.
Our next-generation model: Gemini 1.5
Google. 2024 · 2024
Closest in time.
DeepSpeed-FastGen: High-throughput Text Generation for LLMs via MII and DeepSpeed-Inference. In arXiv
Connor Holmes, Masahiro Tanaka, Michael Wyatt, Ammar Ahmad Awan, Jeff Rasley, Samyam Rajbhandari, Reza Yazdani Aminabadi, Heyang Qin, Arash Bakhtiari, Lev Kurilenko, and Yuxiong He. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Striped Attention: Faster Ring Attention for Causal Transformers. In arXiv
William Brandon, Aniruddha Nrusimha, Kevin Qian, Zachary Ankner, Tian Jin, Zhiye Song, and Jonathan Ragan-Kelley. 2023 · 2023
Cited alongside, same era.
FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning. In arXiv
Tri Dao. 2023 · 2023
Cited alongside, same era.
ElasticFlow: An Elastic Serverless Training Platform for Distributed Deep Learning. In ACM ASPLOS
Diandian Gu, Yihao Zhao, Yinmin Zhong, Yifan Xiong, Zhenhua Han, Peng Cheng, Fan Yang, Gang Huang, Xin Jin, and Xuanzhe Liu. 2023 · 2023
Cited alongside, same era.
Deepspeed ulysses: System optimizations for enabling training of extreme long sequence transformer models. In arXiv
Sam Ade Jacobs, Masahiro Tanaka, Chengming Zhang, Minjia Zhang, Leon Song, Samyam Rajbhandari, and Yuxiong He. 2023 · 2023
Cited alongside, same era.
Sia: Heterogeneity-aware, goodput-optimized ML-cluster scheduling. In ACM SOSP
Suhas Jayaram Subramanya, Daiyaan Arfeen, Shouxu Lin, Aurick Qiao, Zhihao Jia, and Gregory R. Ganger. 2023 · 2023
Cited alongside, same era.
Reducing activation recomputation in large transformer models. In Conference on Machine Learning and Systems
Vijay Anand Korthikanti, Jared Casper, Sangkug Lym, Lawrence McAfee, Michael Andersch, Mohammad Shoeybi, and Bryan Catanzaro. 2023 · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention. In ACM SOSP
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. 2023 · 2023
Cited alongside, same era.
Closest in time.
Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads. In arXiv
Cunchen Hu, Heyang Huang, Liangliang Xu, Xusheng Chen, Jiang Xu, Shuang Chen, Hao Feng, Chenxi Wang, Sa Wang, Yungang Bao, Ninghui Sun, and Yizhou Shan. 2024 · 2024
Closest in time.
Mixtral of Experts. In arXiv
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2024 · 2024
Closest in time.
Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache. In arXiv
Bin Lin, Tao Peng, Chen Zhang, Minmin Sun, Lanbo Li, Hanyu Zhao, Wencong Xiao, Qi Xu, Xiafei Qiu, Shen Li, Zhigang Ji, Yong Li, and Wei Lin. 2024 · 2024
Closest in time.
One Queue Is All You Need: Resolving Head-of-Line Blocking in Large Language Model Serving. In arXiv
Archit Patke, Dhemath Reddy, Saurabh Jha, Haoran Qiu, Christian Pinto, Shengkun Cui, Chandra Narayanaswami, Zbigniew Kalbarczyk, and Ravishankar Iyer. 2024 · 2024
Closest in time.
Linear Attention Sequence Parallelism. In arXiv
Weigao Sun, Zhen Qin, Dong Li, Xuyang Shen, Yu Qiao, and Yiran Zhong. 2024 · 2024
Closest in time.
Introducing Qwen1.5
Qwen Team. 2024 · 2024
Closest in time.
dLoRA: Dynamically Orchestrating Requests and Adapters for LoRA LLM Serving. In USENIX OSDI
Bingyang Wu, Ruidong Zhu, Zili Zhang, Peng Sun, Xuanzhe Liu, and Xin Jin. 2024 · 2024
Closest in time.
Efficient Streaming Language Models with Attention Sinks. In International Conference on Learning Representations (ICLR)
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. 2024 · 2024
Closest in time.
LV-Eval: A Balanced Long-Context Benchmark with 5 Length Levels Up to 256K. In arXiv
Tao Yuan, Xuefei Ning, Dong Zhou, Zhijie Yang, Shiyao Li, Minghui Zhuang, Zheyue Tan, Zhuyu Yao, Dahua Lin, Boxun Li, Guohao Dai, Shengen Yan, and Yu Wang. 2024 · 2024
Closest in time.
H2o: Heavy-hitter oracle for efficient generative inference of large language models. In Neural Information Processing Systems
Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, et al · 2024
Closest in time.
DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving. In USENIX OSDI
Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang. 2024 · 2024
Closest in time.