Attentionstore: Cost-effective attention reuse across multi-turn conversations in large language model serving
Original
Bin Gao, Zhuomin He, Puru Sharma, Qingxuan Kang, Djordje Jevdjic, Junbo Deng, Xingkun Yang, Zhou Yu, and Pengfei Zuo · 2024
Closest in time.
Preble: Efficient distributed prompt scheduling for llm serving
Vikranth Srivatsa, Zijian He, Reyna Abhyankar, Dongming Li, and Yiying Zhang · 2024
Closest in time.
H3dm: A high-bandwidth high-capacity hybrid 3d memory design for gpus
Negar Akbarzadeh, Sina Darabi, Atiyeh Gheibi-Fetrat, Amir Mirzaei, Mohammad Sadrosadati, and Hamid Sarbazi-Azad · 2024
Closest in time.
Efficient and economic large language model inference with attention offloading, 2024
Shaoyuan Chen, Yutong Lin, Mingxing Zhang, and Yongwei Wu · 2024
Closest in time.
Effectively compress kv heads for llm, 2024
Hao Yu, Zelan Yang, Shen Li, Yong Li, and Jianxin Wu · 2024
Closest in time.
Pyramidkv: Dynamic kv cache compression based on pyramidal information funneling, 2024
Zefan Cai., Yichi Zhang, Bofei Gao, Yuliang Liu, Tianyu Liu, Keming Lu, Wayne Xiong, Yue Dong, Baobao Chang, Junjie Hu, and Wen Xiao · 2024
Closest in time.
Zipcache: Accurate and efficient kv cache quantization with salient token identification, 2024
Yefei He, Luoming Zhang, Weijia Wu, Jing Liu, Hong Zhou, and Bohan Zhuang · 2024
Closest in time.
Intactkv: Improving large language model quantization by keeping pivot tokens intact, 2024
Ruikang Liu, Haoli Bai, Haokun Lin, Yuening Li, Han Gao, Zhengzhuo Xu, Lu Hou, Jun Yao, and Chun Yuan · 2024
Closest in time.
Minicache: Kv cache compression in depth dimension for large language models, 2024
Akide Liu, Jing Liu, Zizheng Pan, Yefei He, Gholamreza Haffari, and Bohan Zhuang · 2024
Closest in time.
Attention score is not all you need for token importance indicator in kv cache reduction: Value also matters, 2024
Zhiyu Guo, Hidetaka Kamigaito, and Taro Watanabe · 2024
Closest in time.
A simple and effective l 2 l_{2} norm-based strategy for kv cache compression, 2024
Alessio Devoto, Yu Zhao, Simone Scardapane, and Pasquale Minervini · 2024
Closest in time.
Sirllm: Streaming infinite retentive llm, 2024
Yao Yao, Zuchao Li, and Hai Zhao · 2024
Closest in time.
Pyramidinfer: Pyramid kv cache compression for high-throughput llm inference, 2024
Dongjie Yang, XiaoDong Han, Yan Gao, Yao Hu, Shilin Zhang, and Hai Zhao · 2024
Closest in time.
Keyformer: Kv cache reduction through key tokens selection for efficient generative inference, 2024
Muhammad Adnan, Akhil Arunkumar, Gaurav Jain, Prashant J. Nair, Ilya Soloveychik, and Purushotham Kamath · 2024
Closest in time.
Snapkv: Llm knows what you are looking for before generation, 2024
Yuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh, Acyr Locatelli, Hanchen Ye, Tianle Cai, Patrick Lewis, and Deming Chen · 2024
Closest in time.
Mlkv: Multi-layer key-value heads for memory efficient transformer decoding, 2024
Zayd Muhammad Kawakibi Zuhri, Muhammad Farid Adilazuarda, Ayu Purwarianti, and Alham Fikri Aji · 2024
Closest in time.
Layer-condensed kv cache for efficient inference of large language models, 2024
Haoyi Wu and Kewei Tu · 2024
Closest in time.
You only cache once: Decoder-decoder architectures for language models, 2024
Yutao Sun, Li Dong, Yi Zhu, Shaohan Huang, Wenhui Wang, Shuming Ma, Quanlu Zhang, Jianyong Wang, and Furu Wei · 2024
Closest in time.
Mamba: Linear-time sequence modeling with selective state spaces, 2024
Albert Gu and Tri Dao · 2024
Closest in time.
Transformers are ssms: Generalized models and efficient algorithms through structured state space duality, 2024
Tri Dao and Albert Gu · 2024
Closest in time.
Jamba: A hybrid transformer-mamba language model, 2024
Opher Lieber, Barak Lenz, Hofit Bata, Gal Cohen, Jhonathan Osin, Itay Dalmedigos, Erez Safahi, Shaked Meirom, Yonatan Belinkov, Shai Shalev-Shwartz, Omri Abend, Raz Alon, Tomer Asida, Amir Bergman, Roman Glozman, Michael Gokhman, Avashalom Manevich, Nir Ratner, Noam Rozen, Erez Shwartz, Mor Zusman, and Yoav Shoham · 2024
Closest in time.
Recurrentgemma: Moving past transformers for efficient open language models, 2024
Aleksandar Botev, Soham De, Samuel L Smith, Anushan Fernando, George-Cristian Muraru, Ruba Haroun, Leonard Berrada, Razvan Pascanu, Pier Giuseppe Sessa, Robert Dadashi, Léonard Hussenot, Johan Ferret, Sertan Girgin, Olivier Bachem, Alek Andreev, Kathleen Kenealy, Thomas Mesnard, Cassidy Hardin, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, Pouya Tafti, Armand Joulin, Noah Fiedel, Evan Senter, Yutian Chen, Srivatsan Srinivasan, Guillaume Desjardins, David Budden, Arnaud Doucet, Sharad Vikram, Adam Paszke, Trevor Gale, Sebastian Borgeaud, Charlie Chen, Andy Brock, Antonia Paterson, Jenny Brennan, Meg Risdal, Raj Gundluru, Nesh Devanathan, Paul Mooney, Nilay Chauhan, Phil Culliton, Luiz GUStavo Martins, Elisa Bandy, David Huntsperger, Glenn Cameron, Arthur Zucker, Tris Warkentin, Ludovic Peran, Minh Giang, Zoubin Ghahramani, Clément Farabet, Koray Kavukcuoglu, Demis Hassabis, Raia Hadsell, Yee Whye Teh, and Nando de Frietas · 2024
Closest in time.