Fetching the paper…
Reading the bibliography…
As Large Language Models (LLMs) evolve from text-completion tools into fully fledged agents operating in dynamic environments, they must address the challenge of continually learning and retaining long-term knowledge.
Les troubles de la mémoire accompagnant des lésions hippocampiques bilatérales
Milner, B · 1962
Earlier work this paper cites.
Retrieval time from semantic memory
Collins, A. M. and Quillian, M. R · 1969
Earlier work this paper cites.
Episodic and semantic memory
Tulving, E · 1972
Earlier work this paper cites.
Working memory
Baddeley, A. D. and Hitch, G. J · 1974
Earlier work this paper cites.
The Hippocampus as a Cognitive Map
O’Keefe, J. and Nadel, L · 1978
Earlier work this paper cites.
Preserved learning and retention of pattern-analyzing skill in amnesia: Dissociation of “knowing how” and “knowing that”
Cohen, N. J. and Squire, L. R · 1980
Earlier work this paper cites.
Working Memory
Baddeley, A. D · 1986
Earlier work this paper cites.
Why there are complementary learning systems in the hippocampus and neocortex: Insights from the successes and failures of connectionist models of learning and memory
McClelland, J., McNaughton, B., and O’Reilly, R · 1995
Earlier work this paper cites.
Structure and function of declarative and nondeclarative memory systems
Squire, L. and Zola, S · 1996
Earlier work this paper cites.
Sensory–perceptual episodic memory and its context: Autobiographical memory
Conway, M · 2001
Earlier work this paper cites.
Scaling laws for neural language models, 2020
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2001
Earlier work this paper cites.
Theories of episodic memory
Mayes, A. and Roberts, N · 2001
Earlier work this paper cites.
Episodic memory in primates
Schwartz, B. and Evans, S · 2001
Earlier work this paper cites.
Hippocampal and neocortical contributions to memory: Advances in the complementary learning systems framework
O’Reilly, R. and Norman, K · 2002
Earlier work this paper cites.
Episodic memory in nonhumans: what, and where, is when?
Hampton, R. and Schwartz, B · 2004
Earlier work this paper cites.
Understanding memory through hippocampal remapping
Colgin, L., Moser, E., and Moser, M · 2008
Earlier work this paper cites.
Can we reconcile the declarative memory and spatial navigation views on hippocampal function?
Eichenbaum, H. and Cohen, N · 2014
Earlier work this paper cites.
Graves, A., Wayne, G., and Danihelka, I · 2014
Earlier work this paper cites.
Complementary learning systems
O’Reilly, R., Bhattacharyya, R., Howard, M., and Ketz, N · 2014
Earlier work this paper cites.
Large-scale simple question answering with memory networks, 2015
Bordes, A., Usunier, N., Chopra, S., and Weston, J · 2015
Earlier work this paper cites.
The hippocampus as a cognitive map . . . of social space
Eichenbaum, H · 2015
Earlier work this paper cites.
Sukhbaatar, S., Szlam, A., Weston, J., and Fergus, R · 2015
Earlier work this paper cites.
Hybrid computing using a neural network with dynamic external memory
Graves, A., Wayne, G., Reynolds, M., Harley, T., Danihelka, I., Grabska-Barwińska, A., Colmenarejo, S. G., Grefenstette, E., Ramalho, T., Agapiou, J., et al · 2016
Earlier work this paper cites.
What learning systems do intelligent agents need? complementary learning systems theory updated
Kumaran, D., Hassabis, D., and McClelland, J. L · 2016
Earlier work this paper cites.
The kanerva machine: A generative distributed memory
Wu, Y., Wayne, G., Graves, A., and Lillicrap, T · 2018
Earlier work this paper cites.
Transformer-XL: Attentive language models beyond a fixed-length context
Dai, Z., Yang, Z., Yang, Y., Carbonell, J., Le, Q., and Salakhutdinov, R · 2019
Earlier work this paper cites.
Episodic memory: Neuronal codes for what, where, and when
Sugar, J. and Moser, M · 2019
Earlier work this paper cites.
Transformers are rnns: Fast autoregressive transformers with linear attention
Katharopoulos, A., Vyas, A., Pappas, N., and Fleuret, F · 2020
Earlier work this paper cites.
Generalization through memorization: Nearest neighbor language models
Khandelwal, U., Levy, O., Jurafsky, D., Zettlemoyer, L., and Lewis, M · 2020
Earlier work this paper cites.
Linear attention mechanism: An efficient attention for semantic segmentation
Li, R., Su, J., Duan, C., and Zheng, S · 2020
Earlier work this paper cites.
Editing factual knowledge in language models, 2021
Cao, N. D., Aziz, W., and Titov, I · 2021
Earlier work this paper cites.
Raise a child in large language model: Towards effective and generalizable fine-tuning
Xu, R., Luo, F., Zhang, Z., Tan, C., Chang, B., Huang, S., and Huang, F · 2021
Cited alongside, same era.
Adaptive semiparametric language models
Yogatama, D., de Masson d’Autume, C., and Kong, L · 2021
Cited alongside, same era.
Arani, E., Sarfraz, F., and Zonooz, B · 2022
Cited alongside, same era.
Recurrent memory transformer, 2022
Bulatov, A., Kuratov, Y., and Burtsev, M. S · 2022
Cited alongside, same era.
LoRA: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2022
Cited alongside, same era.
Model tells you what to discard: Adaptive kv cache compression for llms, 2024
Ge, S., Zhang, Y., Liu, L., Zhang, M., Han, J., and Gao, J · 2024
Later among the works it cites.
Goldstein, D., Obeid, F., Alcaide, E., Song, G., and Cheah, E · 2024
Later among the works it cites.
Model editing at scale leads to gradual and catastrophic forgetting, 2024
Gupta, A., Rao, A., and Anumanchipalli, G · 2024
Later among the works it cites.
LM-infinite: Zero-shot extreme length generalization for large language models
Han, C., Wang, Q., Peng, H., Xiong, W., Chen, Y., Ji, H., and Wang, S · 2024
Later among the works it cites.
Kvquant: Towards 10 million context length llm inference with kv cache quantization, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Mitchell, E., Lin, C., Bosselut, A., Manning, C. D., and Finn, C · 2022
Cited alongside, same era.
Learning by distilling context, 2022
Snell, C., Klein, D., and Zhong, R · 2022
Cited alongside, same era.
Dylora: Parameter efficient tuning of pre-trained models using dynamic search-free low rank adaptation, 2022
Valipour, M., Rezagholizadeh, M., Kobyzev, I., and Ghodsi, A · 2022
Cited alongside, same era.
Memformer: A memory-augmented transformer for sequence modeling
Wu, Q., Lan, Z., Qian, K., Gu, J., Geramifard, A., and Yu, Z · 2022
Cited alongside, same era.
The reversal curse: Llms trained on” a is b” fail to learn” b is a”
Berglund, L., Tong, M., Kaufmann, M., Balesni, M., Stickland, A. C., Korbak, T., and Evans, O · 2023
Cited alongside, same era.
Lift yourself up: Retrieval-augmented text generation with self-memory
Cheng, X., Luo, D., Chen, X., Liu, L., Zhao, D., and Yan, R · 2023
Cited alongside, same era.
Enabling large language models to generate text with citations, 2023
Gao, T., Yen, H., Yu, J., and Chen, D · 2023
Cited alongside, same era.
Hooper, C., Kim, S., Mohammadzadeh, H., Mahoney, M. W., Shao, Y. S., Keutzer, K., and Gholami, A · 2024
Later among the works it cites.
The impact of inference acceleration strategies on bias of llms
Kirsten, E., Habernal, I., Nanda, V., and Zafar, M. B · 2024
Later among the works it cites.
Lee, W., Lee, J., Seo, J., and Sim, J · 2024
Later among the works it cites.
Learning, fast and slow: Single- and many-shot learning in the hippocampus
Liao, Z. and Losonczy, A · 2024
Later among the works it cites.
Lin, B., Zhang, C., Peng, T., Zhao, H., Xiao, W., Sun, M., Liu, A., Zhang, Z., Li, L., Qiu, X., Li, S., Ji, Z., Xie, T., Li, Y., and Lin, W · 2024
Later among the works it cites.
Sparser is faster and less is more: Efficient sparse attention for long-range transformers
Lou, C., Jia, Z., Zheng, Z., and Tu, K · 2024
Later among the works it cites.
Meta knowledge for retrieval augmented large language models, 2024
Mombaerts, L., Ding, T., Banerjee, A., Felice, F., Taws, J., and Borogovac, T · 2024
Later among the works it cites.
Dynamic memory compression: Retrofitting llms for accelerated inference, 2024
Nawrot, P., Łańcucki, A., Chochowski, M., Tarjan, D., and Ponti, E. M · 2024
Later among the works it cites.
Memgpt: Towards llms as operating systems, 2024
Packer, C., Wooders, S., Lin, K., Fang, V., Patil, S. G., Stoica, I., and Gonzalez, J. E · 2024
Later among the works it cites.
Anchor-based large language models, 2024
Pang, J., Ye, F., Wong, D. F., He, X., Chen, W., and Wang, L · 2024
Later among the works it cites.
Graph retrieval-augmented generation: A survey, 2024
Peng, B., Zhu, Y., Liu, Y., Bo, X., Shi, H., Hong, C., Zhang, Y., and Tang, S · 2024
Later among the works it cites.
Assessing episodic memory in llms with sequence order recall tasks, 2024
Pink, M., Vo, V. A., Wu, Q., Mu, J., Turek, J. S., Hasson, U., Norman, K. A., Michelmann, S., Huth, A., and Toneva, M · 2024
Later among the works it cites.
Massive editing for large language models via meta learning, 2024
Tan, C., Zhang, G., and Fu, J · 2024
Later among the works it cites.
Razorattention: Efficient kv cache compression through retrieval heads, 2024
Tang, H., Lin, Y., Lin, J., Han, Q., Hong, S., Yao, Y., and Wang, G · 2024
Later among the works it cites.
Reft: Representation finetuning for language models, 2024
Wu, Z., Arora, A., Wang, Z., Geiger, A., Jurafsky, D., Manning, C. D., and Potts, C · 2024
Later among the works it cites.
Infllm: Training-free long-context extrapolation for llms with an efficient context memory, 2024
Xiao, C., Zhang, P., Han, X., Xiao, G., Lin, Y., Zhang, Z., Liu, Z., and Sun, M · 2024
Later among the works it cites.
Long-context language modeling with parallel context encoding, 2024
Yen, H., Gao, T., and Chen, D · 2024
Later among the works it cites.
Lofit: Localized fine-tuning on llm representations, 2024
Yin, F., Ye, X., and Durrett, G · 2024
Later among the works it cites.
In defense of rag in the era of long-context language models, 2024
Yu, T., Xu, A., and Akkiraju, R · 2024
Later among the works it cites.
Wkvquant: Quantizing weight and key/value cache for large language models gains more, 2024
Yue, Y., Yuan, Z., Duanmu, H., Zhou, S., Wu, J., and Nie, L · 2024
Later among the works it cites.
Sglang: Efficient execution of structured language model programs
Zheng, L., Yin, L., Xie, Z., Sun, C., Huang, J., Yu, C. H., Cao, S., Kozyrakis, C., Stoica, I., Gonzalez, J. E., et al · 2024
Later among the works it cites.
Hipporag: Neurobiologically inspired long-term memory for large language models, 2025
Gutiérrez, B. J., Shu, Y., Gu, Y., Yasunaga, M., and Su, Y · 2025
Closest in time.
Memllm: Finetuning llms to use an explicit read-write memory, 2025
Modarressi, A., Köksal, A., Imani, A., Fayyaz, M., and Schütze, H · 2025
Closest in time.
Agent laboratory: Using llm agents as research assistants
Schmidgall, S., Su, Y., Wang, Z., Sun, X., Wu, J., Yu, X., Liu, J., Liu, Z., and Barsoum, E · 2025
Closest in time.
The mamba in the llama: Distilling and accelerating hybrid models, 2025
Wang, J., Paliotta, D., May, A., Rush, A. M., and Dao, T · 2025
Closest in time.
Memorybank: Enhancing large language models with long-term memory
Zhong, W., Guo, L., Gao, Q., Ye, H., and Wang, Y · 2025
Closest in time.