Fetching the paper…
Reading the bibliography…
Key-Value (KV) caching is a common technique to enhance the computational efficiency of Large Language Models (LLMs), but its memory overhead grows rapidly with input length.
Fast transformer decoding: One write-head is all you need, 2019
Noam Shazeer · 1911
Earlier work this paper cites.
Towards ai-complete question answering: A set of prerequisite toy tasks, 2015
Jason Weston, Antoine Bordes, Sumit Chopra, Alexander M. Rush, Bart van Merriënboer, Armand Joulin, and Tomas Mikolov · 2015
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned
Elena Voita, David Talbot, Fedor Moiseev, Rico Sennrich, and Ivan Titov · 2019
Earlier work this paper cites.
Extractive summarization of long documents by combining global and local context
Wen Xiao and Giuseppe Carenini · 2019
Earlier work this paper cites.
Constructing a multi-hop QA dataset for comprehensive evaluation of reasoning steps
Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa · 2020
Earlier work this paper cites.
Compressive transformers for long-range sequence modelling
Jack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier, and Timothy P. Lillicrap · 2020
Earlier work this paper cites.
Layer-wise pruning of transformer attention heads for efficient language modeling, 2021
Kyuhong Shim, Iksoo Choi, Wonyong Sung, and Jungwook Choi · 2021
Earlier work this paper cites.
An empirical survey on long document summarization: Datasets, models, and metrics
Huan Yee Koh, Jiaxin Ju, Ming Liu, and Shirui Pan · 2022
Earlier work this paper cites.
A fast post-training pruning framework for transformers
Woosuk Kwon, Sehoon Kim, Michael W. Mahoney, Joseph Hassoun, Kurt Keutzer, and Amir Gholami · 2022
Earlier work this paper cites.
In-context learning and induction heads, 2022
Catherine Olsson, Nelson Elhage, Neel Nanda, Nicholas Joseph, Nova DasSarma, Tom Henighan, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Scott Johnston, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
Accelerating large language model decoding with speculative sampling, 2023
Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, Laurent Sifre, and John Jumper · 2023
Cited alongside, same era.
LLMLingua: Compressing prompts for accelerated inference of large language models
Huiqiang Jiang, Qianhui Wu, Chin-Yew Lin, Yuqing Yang, and Lili Qiu · 2023
Cited alongside, same era.
Inference-time intervention: Eliciting truthful answers from a language model
Kenneth Li, Oam Patel, Fernanda Viégas, Hanspeter Pfister, and Martin Wattenberg · 2023
Cited alongside, same era.
A critical evaluation of evaluations for long-form question answering, 2023
Fangyuan Xu, Yixiao Song, Mohit Iyyer, and Eunsol Choi · 2023
Cited alongside, same era.
H2o: Heavy-hitter oracle for efficient generative inference of large language models
Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Re, Clark Barrett, Zhangyang Wang, and Beidi Chen · 2023
Zhuoran Jin, Pengfei Cao, Hongbang Yuan, Yubo Chen, Jiexin Xu, Huaijun Li, Xiaojian Jiang, Kang Liu, and Jun Zhao · 2024
Closest in time.
Babilong: Testing the limits of llms with long context reasoning-in-a-haystack, 2024
Yuri Kuratov, Aydar Bulatov, Petr Anokhin, Ivan Rodkin, Dmitry Sorokin, Artyom Sorokin, and Mikhail Burtsev · 2024
Closest in time.
Dodo: Dynamic contextual compression for decoder-only LMs
Guanghui Qin, Corby Rosset, Ethan Chau, Nikhil Rao, and Benjamin Van Durme · 2024
Closest in time.
Razorattention: Efficient kv cache compression through retrieval heads, 2024
Hanlin Tang, Yang Lin, Jing Lin, Qingsen Han, Shikuan Hong, Yiwu Yao, and Gongyi Wang · 2024
Closest in time.
Retrieval head mechanistically explains long-context factuality, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The claude 3 model family: Opus, sonnet, haiku, March 2024
Anthropic · 2024
Cited alongside, same era.
LongBench: A bilingual, multitask benchmark for long context understanding
Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, Yuxiao Dong, Jie Tang, and Juanzi Li · 2024
Cited alongside, same era.
Pyramidkv: Dynamic kv cache compression based on pyramidal information funneling, 2024
Zefan Cai, Yichi Zhang, Bofei Gao, Yuliang Liu, Tianyu Liu, Keming Lu, Wayne Xiong, Yue Dong, Baobao Chang, Junjie Hu, and Wen Xiao · 2024
Cited alongside, same era.
FlashAttention-2: Faster attention with better parallelism and work partitioning
Tri Dao · 2024
Cited alongside, same era.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Cited alongside, same era.
Ada-kv: Optimizing kv cache eviction by adaptive budget allocation for efficient llm inference, 2024
Yuan Feng, Junlin Lv, Yukun Cao, Xike Xie, and S. Kevin Zhou · 2024
Cited alongside, same era.
LongLLMLingua: Accelerating and enhancing LLMs in long context scenarios via prompt compression
Huiqiang Jiang, Qianhui Wu, Xufang Luo, Dongsheng Li, Chin-Yew Lin, Yuqing Yang, and Lili Qiu · 2024
Cited alongside, same era.
Wenhao Wu, Yizhong Wang, Guangxuan Xiao, Hao Peng, and Yao Fu · 2024
Closest in time.
Efficient streaming language models with attention sinks
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis · 2024
Closest in time.
An Yang, Baosong Yang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Zhou, Chengpeng Li, Chengyuan Li, Dayiheng Liu, Fei Huang, et al · 2024
Closest in time.
Chunkattention: Efficient self-attention with prefix-aware kv cache and two-phase partition, 2024
Lu Ye, Ze Tao, Yong Huang, and Yang Li · 2024
Closest in time.
A survey on recent advances in llm-based multi-turn dialogue systems, 2024
Zihao Yi, Jiarui Ouyang, Yuwen Liu, Tianhao Liao, Zhe Xu, and Ying Shen · 2024
Closest in time.
Attention heads of large language models: A survey, 2024
Zifan Zheng, Yezhaohui Wang, Yuxin Huang, Shichao Song, Mingchuan Yang, Bo Tang, Feiyu Xiong, and Zhiyu Li · 2024
Closest in time.
A survey on efficient inference for large language models, 2024
Zixuan Zhou, Xuefei Ning, Ke Hong, Tianyu Fu, Jiaming Xu, Shiyao Li, Yuming Lou, Luning Wang, Zhihang Yuan, Xiuhong Li, Shengen Yan, Guohao Dai, Xiao-Ping Zhang, Yuhan Dong, and Yu Wang · 2024
Closest in time.