Fetching the paper…
Reading the bibliography…
Large Language Models have excelled in various domains but face efficiency challenges due to the growing Key-Value (KV) cache required for long-sequence inference.
Learning question classifiers
Xin Li and Dan Roth · 2002
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension, 2017
Mandar Joshi, Eunsol Choi, Daniel S. Weld, and Luke Zettlemoyer · 2017
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer · 2017
Earlier work this paper cites.
The narrativeqa reading comprehension challenge
Tomáš Kočiskỳ, Jonathan Schwarz, Phil Blunsom, Chris Dyer, Karl Moritz Hermann, Gábor Melis, and Edward Grefenstette · 2018
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W Cohen, Ruslan Salakhutdinov, and Christopher D Manning · 2018
Earlier work this paper cites.
Multi-news: A large-scale multi-document summarization dataset and abstractive hierarchical model
Alexander R Fabbri, Irene Li, Tianwei She, Suyi Li, and Dragomir R Radev · 2019
Earlier work this paper cites.
Samsum corpus: A human-annotated dialogue dataset for abstractive summarization
Bogdan Gliwa, Iwona Mochol, Maciej Biesek, and Aleksander Wawer · 2019
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan · 2020
Earlier work this paper cites.
Constructing a multi-hop QA dataset for comprehensive evaluation of reasoning steps
Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa · 2020
Earlier work this paper cites.
Constructing a multi-hop qa dataset for comprehensive evaluation of reasoning steps
Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa · 2020
Earlier work this paper cites.
A dataset of information-seeking questions and answers anchored in research papers
Pradeep Dasigi, Kyle Lo, Iz Beltagy, Arman Cohan, Noah A Smith, and Matt Gardner · 2021
Earlier work this paper cites.
Efficient attentions for long document summarization
Luyang Huang, Shuyang Cao, Nikolaus Parulian, Heng Ji, and Lu Wang · 2021
Earlier work this paper cites.
Qmsum: A new benchmark for query-based multi-domain meeting summarization
Ming Zhong, Da Yin, Tao Yu, Ahmad Zaidi, Mutethia Mutuma, Rahul Jha, Ahmed Hassan Awadallah, Asli Celikyilmaz, Yang Liu, Xipeng Qiu, et al · 2021
Earlier work this paper cites.
Flashattention: Fast and memory-efficient exact attention with io-awareness
Tri Dao, Dan Fu, Stefano Ermon, Atri Rudra, and Christopher Ré · 2022
Earlier work this paper cites.
Zeroquant: Efficient and affordable post-training quantization for large-scale transformers
Zhewei Yao, Reza Yazdani Aminabadi, Minjia Zhang, Xiaoxia Wu, Conglong Li, and Yuxiong He · 2022
Earlier work this paper cites.
Musique: Multihop questions via single-hop question composition
Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal · 2022
Earlier work this paper cites.
Summedits: measuring llm ability at factual reasoning through the lens of summarization
Philippe Laban, Wojciech Kryściński, Divyansh Agarwal, Alexander Richard Fabbri, Caiming Xiong, Shafiq Joty, and Chien-Sheng Wu · 2023
Earlier work this paper cites.
Llm-based code generation method for golang compiler testing
Qiuhan Gu · 2023
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
Model tells you what to discard: Adaptive kv cache compression for llms
Suyu Ge, Yunan Zhang, Liyuan Liu, Minjia Zhang, Jiawei Han, and Jianfeng Gao · 2023
Earlier work this paper cites.
Catalyst: Optimizing cache management for large in-memory key-value systems
Kefei Wang and Feng Chen · 2023
Earlier work this paper cites.
Efficient streaming language models with attention sinks
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis · 2023
Earlier work this paper cites.
Deja vu: Contextual sparsity for efficient llms at inference time, 2023
Zichang Liu, Jue Wang, Tri Dao, Tianyi Zhou, Binhang Yuan, Zhao Song, Anshumali Shrivastava, Ce Zhang, Yuandong Tian, Christopher Re, and Beidi Chen · 2023
Earlier work this paper cites.
Flashattention-2: Faster attention with better parallelism and work partitioning
Tri Dao · 2023
Cited alongside, same era.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica · 2023
Cited alongside, same era.
Albert Q Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, et al · 2023
Cited alongside, same era.
Gqa: Training generalized multi-query transformer models from multi-head checkpoints, 2023
Joshua Ainslie, James Lee-Thorp, Michiel de Jong, Yury Zemlyanskiy, Federico Lebrón, and Sumit Sanghai · 2023
Cited alongside, same era.
Longbench: A bilingual, multitask benchmark for long context understanding
Minference 1.0: Accelerating pre-filling for long-context llms via dynamic sparse attention
Huiqiang Jiang, Yucheng Li, Chengruidong Zhang, Qianhui Wu, Xufang Luo, Surin Ahn, Zhenhua Han, Amir H Abdi, Dongsheng Li, Chin-Yew Lin, et al · 2024
Closest in time.
Retrievalattention: Accelerating long-context llm inference via vector retrieval
Di Liu, Meng Chen, Baotong Lu, Huiqiang Jiang, Zhenhua Han, Qianxi Zhang, Qi Chen, Chengruidong Zhang, Bailu Ding, Kai Zhang, et al · 2024
Closest in time.
Arkvale: Efficient generative llm inference with recallable key-value eviction
Renze Chen, Zhuofeng Wang, Beiquan Cao, Tong Wu, Size Zheng, Xiuhong Li, Xuechao Wei, Shengen Yan, Meng Li, and Yun Liang · 2024
Closest in time.
The llama 3 herd of models, 2024
Aaron Grattafiori, Abhimanyu Dubey, and Abhinav Jauhri .et al · 2024
Closest in time.
Ruler: What’s the real context size of your long-context language models?
Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, and Boris Ginsburg · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, et al · 2023
Cited alongside, same era.
Needle In A Haystack - pressure testing LLMs
Gregory Kamradt · 2023
Cited alongside, same era.
Longcoder: A long-range pre-trained language model for code completion, 2023
Daya Guo, Canwen Xu, Nan Duan, Jian Yin, and Julian McAuley · 2023
Cited alongside, same era.
Repobench: Benchmarking repository-level code auto-completion systems, 2023
Tianyang Liu, Canwen Xu, and Julian McAuley · 2023
Cited alongside, same era.
Lost in the middle: How language models use long contexts
Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang · 2023
Cited alongside, same era.
Longcoder: A long-range pre-trained language model for code completion
Daya Guo, Canwen Xu, Nan Duan, Jian Yin, and Julian McAuley · 2023
Cited alongside, same era.
Repobench: Benchmarking repository-level code auto-completion systems
Tianyang Liu, Canwen Xu, and Julian McAuley · 2023
Cited alongside, same era.
A survey on recent advances in llm-based multi-turn dialogue systems
Zihao Yi, Jiarui Ouyang, Yuwen Liu, Tianhao Liao, Zhe Xu, and Ying Shen · 2024
Cited alongside, same era.
Closest in time.
Prompt cache: Modular attention reuse for low-latency inference
In Gim, Guojun Chen, Seung-seob Lee, Nikhil Sarda, Anurag Khandelwal, and Lin Zhong · 2024
Closest in time.
Sglang: Efficient execution of structured language model programs
Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E Gonzalez, et al · 2024
Closest in time.
Duoattention: Efficient long-context llm inference with retrieval and streaming heads
Guangxuan Xiao, Jiaming Tang, Jingwei Zuo, Junxian Guo, Shang Yang, Haotian Tang, Yao Fu, and Song Han · 2024
Closest in time.
Yu Fu, Zefan Cai, Abedelkadir Asi, Wayne Xiong, Yue Dong, and Wen Xiao · 2024
Closest in time.
Kivi: A tuning-free asymmetric 2bit quantization for kv cache
Zirui Liu, Jiayi Yuan, Hongye Jin, Shaochen Zhong, Zhaozhuo Xu, Vladimir Braverman, Beidi Chen, and Xia Hu · 2024
Closest in time.
Qaq: Quality adaptive quantization for llm kv cache
Shichen Dong, Wen Cheng, Jiayu Qin, and Wei Wang · 2024
Closest in time.
Triforce: Lossless acceleration of long sequence generation with hierarchical speculative decoding
Hanshi Sun, Zhuoming Chen, Xinyu Yang, Yuandong Tian, and Beidi Chen · 2024
Closest in time.
Zoology: Measuring and improving recall in efficient language models
Simran Arora, Sabri Eyuboglu, Aman Timalsina, Isys Johnson, Michael Poli, James Zou, Atri Rudra, and Christopher Ré · 2024
Closest in time.
Yucheng Li, Huiqiang Jiang, Chengruidong Zhang, Qianhui Wu, Xufang Luo, Surin Ahn, Amir H Abdi, Dongsheng Li, Jianfeng Gao, Yuqing Yang, et al · 2025
Closest in time.
Pqcache: Product quantization-based kvcache for long context llm inference
Hailin Zhang, Xiaodong Ji, Yilin Chen, Fangcheng Fu, Xupeng Miao, Xiaonan Nie, Weipeng Chen, and Bin Cui · 2025
Closest in time.
He Sun, Li Li, Mingjun Xiao, and Chengzhong Xu · 2025
Closest in time.
Kvzip: Query-agnostic kv cache compression with context reconstruction
Jang-Hyun Kim, Jinuk Kim, Sangwoo Kwon, Jae W Lee, Sangdoo Yun, and Hyun Oh Song · 2025
Closest in time.
Expected attention: Kv cache compression by estimating attention from future queries distribution
Alessio Devoto, Maximilian Jeblick, and Simon Jégou · 2025
Closest in time.
Draft-based approximate inference for llms
Kevin Galim, Ethan Ewer, Wonjun Kang, Minjae Lee, Hyung Il Koo, and Kangwook Lee · 2025
Closest in time.
Identify critical kv cache in llm inference from an output perturbation perspective
Yuan Feng, Junlin Lv, Yukun Cao, Xike Xie, and S Kevin Zhou · 2025
Closest in time.
Taming the fragility of kv cache eviction in llm inference, 2025
Yuan Feng, Haoyu Guo, JunLin Lv, S. Kevin Zhou, and Xike Xie · 2025
Closest in time.
Specvlm: Enhancing speculative decoding of video llms via verifier-guided token pruning
Yicheng Ji, Jun Zhang, Heming Xia, Jinpeng Chen, Lidan Shou, Gang Chen, and Huan Li · 2025
Closest in time.