Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have become more prevalent in long-context applications such as interactive chatbots, document analysis, and agent workflows, but it is challenging to serve long-context requests with low latency and high throughput.
Compressive transformers for long-range sequence modelling, 2019
Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, and Timothy P. Lillicrap · 1911
Earlier work this paper cites.
Retrieval-augmented generation for knowledge-intensive nlp tasks, 2021
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen tau Yih, Tim Rocktäschel, Sebastian Riedel, and Douwe Kiela · 2005
Earlier work this paper cites.
Evaluating large language models trained on code, 2021
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Josh Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba · 2021
Earlier work this paper cites.
Memory-efficient transformers via top- k k attention, 2021
Ankit Gupta, Guy Dar, Shaya Goodman, David Ciprut, and Jonathan Berant · 2021
Earlier work this paper cites.
Deepspeed inference: Enabling efficient inference of transformer models at unprecedented scale, 2022
Reza Yazdani Aminabadi, Samyam Rajbhandari, Minjia Zhang, Ammar Ahmad Awan, Cheng Li, Du Li, Elton Zheng, Jeff Rasley, Shaden Smith, Olatunji Ruwase, and Yuxiong He · 2022
Earlier work this paper cites.
Fast inference from transformers via speculative decoding
Yaniv Leviathan, Matan Kalman, and Yossi Matias · 2022
Earlier work this paper cites.
Orca: A distributed serving system for Transformer-Based generative models
Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
Accelerating large language model decoding with speculative sampling
Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, Laurent Sifre, and John Jumper · 2023
Earlier work this paper cites.
Flashattention-2: Faster attention with better parallelism and work partitioning, 2023
Tri Dao · 2023
Earlier work this paper cites.
Gptq: Accurate post-training quantization for generative pre-trained transformers, 2023
Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh · 2023
Earlier work this paper cites.
Flashdecoding++: Faster large language model inference on gpus
Ke Hong, Guohao Dai, Jiaming Xu, Qiuli Mao, Xiuhong Li, Jun Liu, Kangdi Chen, Hanyu Dong, and Yu Wang · 2023
Earlier work this paper cites.
vllm: Easy, fast, and cheap llm serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Yu, Joseph E Gonzalez, Hao Zhang, and Ion Stoica · 2023
Earlier work this paper cites.
Visual instruction tuning, 2023
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee · 2023
Cited alongside, same era.
Llm-pruner: On the structural pruning of large language models, 2023
Xinyin Ma, Gongfan Fang, and Xinchao Wang · 2023
Cited alongside, same era.
Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng, Zeyu Wang, Rae Ying Yee Wong, Zhuoming Chen, Daiyaan Arfeen, Reyna Abhyankar, and Zhihao Jia · 2023
Cited alongside, same era.
Gpt-fast, 2023
pytorch-labs · 2023
Cited alongside, same era.
The synergy of speculative decoding and batching in serving large language models, 2023
Qidong Su, Christina Giannoula, and Gennady Pekhimenko · 2023
Ruler: What’s the real context size of your long-context language models?, 2024
Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, and Boris Ginsburg · 2024
Closest in time.
Snapkv: Llm knows what you are looking for before generation, 2024
Yuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh, Acyr Locatelli, Hanchen Ye, Tianle Cai, Patrick Lewis, and Deming Chen · 2024
Closest in time.
Transformers are multi-state rnns, 2024
Matanel Oren, Michael Hassid, Nir Yarden, Yossi Adi, and Roy Schwartz · 2024
Closest in time.
vattention: Dynamic memory management for serving llms without pagedattention, 2024
Ramya Prabhu, Ajay Nayak, Jayashree Mohan, Ramachandran Ramjee, and Ashish Panwar · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
MLC-LLM, 2023
MLC team · 2023
Cited alongside, same era.
Speculative decoding: Exploiting speculative execution for accelerating seq2seq generation, 2023
Heming Xia, Tao Ge, Peiyi Wang, Si-Qing Chen, Furu Wei, and Zhifang Sui · 2023
Cited alongside, same era.
H 2 o: Heavy-hitter oracle for efficient generative inference of large language models, 2023
Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, Zhangyang Wang, and Beidi Chen · 2023
Cited alongside, same era.
The llama 3 herd of models, 2024
AI@Meta · 2024
Cited alongside, same era.
Codewhisperer, 2024
Amazon AWS · 2024
Cited alongside, same era.
Pyramidkv: Dynamic kv cache compression based on pyramidal information funneling, 2024
Zefan Cai., Yichi Zhang, Bofei Gao, Yuliang Liu, Tianyu Liu, Keming Lu, Wayne Xiong, Yue Dong, Baobao Chang, Junjie Hu, and Wen Xiao · 2024
Cited alongside, same era.
Our next-generation model: Gemini 1.5, 2024
Google Deepmind · 2024
Cited alongside, same era.
QwenTeam · 2024
Closest in time.
Loki: Low-rank keys for efficient sparse attention, 2024
Prajwal Singhania, Siddharth Singh, Shwai He, Soheil Feizi, and Abhinav Bhatele · 2024
Closest in time.
Quest: Query-aware sparsity for efficient long-context llm inference, 2024
Jiaming Tang, Yilong Zhao, Kan Zhu, Guangxuan Xiao, Baris Kasikci, and Song Han · 2024
Closest in time.
Distributed speculative inference of large language models is provably faster, 2024
Nadav Timor, Jonathan Mamou, Daniel Korat, Moshe Berchansky, Oren Pereg, Moshe Wasserblat, Tomer Galanti, Michal Gordon, and David Harel · 2024
Closest in time.
Pyramidinfer: Pyramid kv cache compression for high-throughput llm inference, 2024
Dongjie Yang, XiaoDong Han, Yan Gao, Yao Hu, Shilin Zhang, and Hai Zhao · 2024
Closest in time.
Llm inference unveiled: Survey and roofline model insights, 2024
Zhihang Yuan, Yuzhang Shang, Yang Zhou, Zhen Dong, Chenhao Xue, Bingzhe Wu, Zhikai Li, Qingyi Gu, Yong Jae Lee, Yan Yan, Beidi Chen, Guangyu Sun, and Kurt Keutzer · 2024
Closest in time.
Pqcache: Product quantization-based kvcache for long context llm inference, 2024
Hailin Zhang, Xiaodong Ji, Yilin Chen, Fangcheng Fu, Xupeng Miao, Xiaonan Nie, Weipeng Chen, and Bin Cui · 2024
Closest in time.
Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang · 2024
Closest in time.