Fetching the paper…
Reading the bibliography…
This paper examines memory mechanisms in Large Language Models (LLMs), emphasizing their importance for context-rich responses, reduced hallucinations, and improved efficiency.
A unified cognitive architecture for physical agents
Pat Langley and Dongkyu Choi · 2006
Earlier work this paper cites.
The stanford corenlp natural language processing toolkit
Christopher D Manning, Mihai Surdeanu, John Bauer, Jenny Rose Finkel, Steven Bethard, and David McClosky · 2014
Earlier work this paper cites.
The Soar cognitive architecture
John E Laird · 2019
Earlier work this paper cites.
Act-r: A cognitive architecture for modeling cognition
Frank E Ritter, Farnaz Tehranchi, and Jacob D Oury · 2019
Earlier work this paper cites.
Etc: Encoding long and structured inputs in transformers
Joshua Ainslie, Santiago Ontanon, Chris Alberti, Vaclav Cvicek, Zachary Fisher, Philip Pham, Anirudh Ravula, Sumit Sanghai, Qifan Wang, and Li Yang · 2020
Earlier work this paper cites.
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan · 2020
Earlier work this paper cites.
Reformer: The efficient transformer
Nikita Kitaev, Łukasz Kaiser, and Anselm Levskaya · 2020
Earlier work this paper cites.
40 years of cognitive architectures: core cognitive abilities and practical applications
Iuliia Kotseruba and John K Tsotsos · 2020
Earlier work this paper cites.
Test-time training with self-supervision for generalization under distribution shifts
Yu Sun, Xiaolong Wang, Zhuang Liu, John Miller, Alexei Efros, and Moritz Hardt · 2020
Earlier work this paper cites.
Efficient content-based sparse attention with routing transformers
Aurko Roy, Mohammad Saffar, Ashish Vaswani, and David Grangier · 2021
Earlier work this paper cites.
Test-time training with masked autoencoders
Yossi Gandelsman, Yu Sun, Xinlei Chen, and Alexei Efros · 2022
Earlier work this paper cites.
Longt5: Efficient text-to-text transformer for long sequences
Mandy Guo, Joshua Ainslie, David Uthus, Santiago Ontanon, Jianmo Ni, Yun-Hsuan Sung, and Yinfei Yang · 2022
Earlier work this paper cites.
Learned token pruning for transformers
Sehoon Kim, Sheng Shen, David Thorsley, Amir Gholami, Woosuk Kwon, Joseph Hassoun, and Kurt Keutzer · 2022
Earlier work this paper cites.
Sparse token transformer with attention back tracking
Heejun Lee, Minki Kang, Youngwan Lee, and Sung Ju Hwang · 2022
Earlier work this paper cites.
Relational memory-augmented language models
Qi Liu, Dani Yogatama, and Phil Blunsom · 2022
Earlier work this paper cites.
Dynamic context pruning for efficient and interpretable autoregressive transformers
Sotiris Anagnostidis, Dario Pavllo, Luca Biggio, Lorenzo Noci, Aurelien Lucchi, and Thomas Hofmann · 2023
Earlier work this paper cites.
Max-margin token selection in attention mechanism
Davoud Ataee Tarzanagh, Yingcong Li, Xuechen Zhang, and Samet Oymak · 2023
Earlier work this paper cites.
Unlimiformer: Long-range transformers with unlimited length input
Amanda Bertsch, Uri Alon, Graham Neubig, and Matthew Gormley · 2023
Earlier work this paper cites.
Adapting language models to compress contexts
Alexis Chevalier, Alexander Wettig, Anirudh Ajith, and Danqi Chen · 2023
Earlier work this paper cites.
Longnet: Scaling transformers to 1,000,000,000 tokens
Jiayu Ding, Shuming Ma, Li Dong, Xingxing Zhang, Shaohan Huang, Wenhui Wang, Nanning Zheng, and Furu Wei · 2023
Earlier work this paper cites.
Empower your model with longer and better context comprehension
Yifei Gao, Lei Wang, Jun Fang, Longhua Hu, and Jun Cheng · 2023
Earlier work this paper cites.
Empowering working memory for large language model agents
Jing Guo, Nan Li, Jianchuan Qi, Hang Yang, Ruiqiao Li, Yuzhen Feng, Si Zhang, and Ming Xu · 2023
Earlier work this paper cites.
Test-time training on nearest neighbors for large language models
Moritz Hardt and Yu Sun · 2023
Earlier work this paper cites.
Memory matters: The need to improve long-term memory in llm-agents
Kostas Hatalis, Despina Christou, Joshua Myers, Steven Jones, Keith Lambert, Adam Amos-Binks, Zohreh Dannenhauer, and Dustin Dannenhauer · 2023
Earlier work this paper cites.
Efficient long-text understanding with short-text models
Maor Ivgi, Uri Shaham, and Jonathan Berant · 2023
Earlier work this paper cites.
Compressed context memory for online language model interaction
Jang-Hyun Kim, Junyoung Yeom, Sangdoo Yun, and Hyun Oh Song · 2023
Earlier work this paper cites.
Efficient memory management for large language model serving with pagedattention
Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica · 2023
Earlier work this paper cites.
Think-in-memory: Recalling and post-thinking enable llms with long-term memory
Lei Liu, Xiaoyan Yang, Yue Shen, Binbin Hu, Zhiqiang Zhang, Jinjie Gu, and Guannan Zhang · 2023
Cited alongside, same era.
Landmark attention: Random-access infinite context length for transformers
Amirkeivan Mohtashami and Martin Jaggi · 2023
Cited alongside, same era.
Faster causal attention over large sequences through sparse flash attention
Matteo Pagliardini, Daniele Paliotta, Martin Jaggi, and François Fleuret · 2023
Cited alongside, same era.
Rwkv: Reinventing rnns for the transformer era
Bo Peng, Eric Alcaide, Quentin Anthony, Alon Albalak, Samuel Arcadinho, Stella Biderman, Huanqi Cao, Xin Cheng, Michael Chung, Matteo Grella, et al · 2023
Cited alongside, same era.
Efficient llm inference with kcache
Qiaozhi He and Zhihua Wu · 2024
Later among the works it cites.
Hmt: Hierarchical memory transformer for long context language processing
Zifan He, Zongyue Qin, Neha Prakriya, Yizhou Sun, and Jason Cong · 2024
Later among the works it cites.
Memserve: Context caching for disaggregated llm serving with elastic memory pool
Cunchen Hu, Heyang Huang, Junhao Hu, Jiang Xu, Xusheng Chen, Tao Xie, Chenxi Wang, Sa Wang, Yungang Bao, Ninghui Sun, et al · 2024
Later among the works it cites.
Transformerfam: Feedback attention is working memory
Dongseong Hwang, Weiran Wang, Zhuoyuan Huo, Khe Chai Sim, and Pedro Moreno Mengibar · 2024
Later among the works it cites.
A2sf: Accumulative attention scoring with forgetting factor for token pruning in transformer decoder
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Robert Praas · 2023
Cited alongside, same era.
Zebra: Extending context window with layerwise grouped local-global attention
Kaiqiang Song, Xiaoyang Wang, Sangwoo Cho, Xiaoman Pan, and Dong Yu · 2023
Cited alongside, same era.
Recursively summarizing enables long-term dialogue memory in large language models
Qingyue Wang, Liang Ding, Yanan Cao, Zhiliang Tian, Shi Wang, Dacheng Tao, and Li Guo · 2023
Cited alongside, same era.
Efficient streaming language models with attention sinks
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis · 2023
Cited alongside, same era.
Chunk, align, select: A simple long-sequence processing method for transformers
Jiawen Xie, Pengyu Cheng, Xiao Liang, Yong Dai, and Nan Du · 2023
Cited alongside, same era.
Megabyte: Predicting million-byte sequences with multiscale transformers
Lili Yu, Dániel Simig, Colin Flaherty, Armen Aghajanyan, Luke Zettlemoyer, and Mike Lewis · 2023
Cited alongside, same era.
Evolving large language model assistant with long-term conditional memory
Ruifeng Yuan, Shichao Sun, Zili Wang, Ziqiang Cao, and Wenjie Li · 2023
Cited alongside, same era.
H2o: Heavy-hitter oracle for efficient generative inference of large language models
Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark Barrett, et al · 2023
Cited alongside, same era.
Hyun-rae Jo and Dongkun Shin · 2024
Later among the works it cites.
Hana Kim, Kai Tzu-iunn Ong, Seoyeon Kim, Dongha Lee, and Jinyoung Yeo · 2024
Later among the works it cites.
A human-inspired reading agent with gist memory of very long contexts
Kuang-Huei Lee, Xinyun Chen, Hiroki Furuta, John Canny, and Ian Fischer · 2024
Later among the works it cites.
Snapkv: Llm knows what you are looking for before generation
Yuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh, Acyr Locatelli, Hanchen Ye, Tianle Cai, Patrick Lewis, and Deming Chen · 2024
Later among the works it cites.
Beyond kv caching: Shared attention for efficient llms
Bingli Liao and Danilo Vasconcellos Vargas · 2024
Later among the works it cites.
Infinite-llm: Efficient llm service for long context with distattention and distributed kvcache
Bin Lin, Chen Zhang, Tao Peng, Hanyu Zhao, Wencong Xiao, Minmin Sun, Anmin Liu, Zhipeng Zhang, Lanbo Li, Xiafei Qiu, et al · 2024
Later among the works it cites.
Longheads: Multi-head attention is secretly a long context processor
Yi Lu, Xin Zhou, Wei He, Jun Zhao, Tao Ji, Tao Gui, Qi Zhang, and Xuanjing Huang · 2024
Later among the works it cites.
Taking a deep breath: Enhancing language modeling of large language models with sentinel tokens
Weiyao Luo, Suncong Zheng, Heming Xia, Weikang Wang, Yan Lei, Tianyu Liu, Shuang Chen, and Zhifang Sui · 2024
Later among the works it cites.
Cross-layer attention sharing for large language models
Yongyu Mu, Yuzhang Wu, Yuchun Fan, Chenglong Wang, Hengyu Li, Qiaozhi He, Murun Yang, Tong Xiao, and Jingbo Zhu · 2024
Later among the works it cites.
Dynamic memory compression: Retrofitting llms for accelerated inference
Piotr Nawrot, Adrian Łańcucki, Marcin Chochowski, David Tarjan, and Edoardo M Ponti · 2024
Later among the works it cites.
Instinfer: In-storage attention offloading for cost-effective long-context llm inference
Xiurui Pan, Endian Li, Qiao Li, Shengwen Liang, Yizhou Shan, Ke Zhou, Yingwei Luo, Xiaolin Wang, and Jie Zhang · 2024
Later among the works it cites.
vattention: Dynamic memory management for serving llms without pagedattention
Ramya Prabhu, Ajay Nayak, Jayashree Mohan, Ramachandran Ramjee, and Ashish Panwar · 2024
Later among the works it cites.
Branch-train-mix: Mixing expert llms into a mixture-of-experts llm
Sainbayar Sukhbaatar, Olga Golovneva, Vasu Sharma, Hu Xu, Xi Victoria Lin, Baptiste Rozière, Jacob Kahn, Daniel Li, Wen-tau Yih, Jason Weston, et al · 2024
Later among the works it cites.
D2o: Dynamic discriminative operations for efficient generative inference of large language models
Zhongwei Wan, Xinjian Wu, Yu Zhang, Yi Xin, Chaofan Tao, Zhihong Zhu, Xin Wang, Siqi Luo, Jing Xiong, and Mi Zhang · 2024
Later among the works it cites.
Training-free exponential extension of sliding window context with cascading kv cache
Jeffrey Willette, Heejun Lee, Youngwan Lee, Myeongjae Jeon, and Sung Ju Hwang · 2024
Later among the works it cites.
Think: Thinner key cache by query-driven pruning
Yuhui Xu, Zhanming Jie, Hanze Dong, Lei Wang, Xudong Lu, Aojun Zhou, Amrita Saha, Caiming Xiong, and Doyen Sahoo · 2024
Later among the works it cites.
Chunkattention: Efficient self-attention with prefix-aware kv cache and two-phase partition
Lu Ye, Ze Tao, Yong Huang, and Yang Li · 2024
Later among the works it cites.
Effectively compress kv heads for llm
Hao Yu, Zelan Yang, Shen Li, Yong Li, and Jianxin Wu · 2024
Later among the works it cites.
Efficient sparse attention needs adaptive token release
Chaoran Zhang, Lixin Zou, Dan Luo, Min Tang, Xiangyang Luo, Zihao Li, and Chenliang Li · 2024
Later among the works it cites.
Alisa: Accelerating large language model inference via sparsity-aware kv caching
Youpeng Zhao, Di Wu, and Jun Wang · 2024
Later among the works it cites.
Memorybank: Enhancing large language models with long-term memory
Wanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye, and Yanlin Wang · 2024
Later among the works it cites.
Mlkv: Multi-layer key-value heads for memory efficient transformer decoding
Zayd Muhammad Kawakibi Zuhri, Muhammad Farid Adilazuarda, Ayu Purwarianti, and Alham Fikri Aji · 2024
Later among the works it cites.