Fetching the paper…
Reading the bibliography…
As new knowledge rapidly accumulates, language models (LMs) with pretrained knowledge quickly become obsolete.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J. Cohen · 1989
Earlier work this paper cites.
Connectionist models of recognition memory: Constraints imposed by learning and forgetting functions
R. Ratcliff · 1990
Earlier work this paper cites.
Catastrophic forgetting, rehearsal and pseudorehearsal
Anthony Robins · 1995
Earlier work this paper cites.
Natural Language Processing with Python
Steven Bird, Edward Loper, and Ewan Klein · 2009
Earlier work this paper cites.
Semantic parsing on Freebase from question-answer pairs
Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang · 2013
Earlier work this paper cites.
Playing Atari with deep reinforcement learning
Volodymyr Mnih, Koray Kavukcuoglu, David Silver, Alex Graves Ioannis Antonoglou, Daan Wierstra, and Martin Riedmiller · 2013
Earlier work this paper cites.
TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension
Mandar Joshi, Eunsol Choi, Daniel S Weld, and Luke Zettlemoyer · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A. Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, Demis Hassabis, Claudia Clopath, Dharshan Kumaran, and Raia Hadsell · 2017
Earlier work this paper cites.
Learning without forgetting
Zhizhong Li and Derek Hoiem · 2017
Earlier work this paper cites.
Gradient episodic memory for continual learning
David Lopez-Paz and Marc’Aurelio Ranzato · 2017
Earlier work this paper cites.
Continual learning with deep generative replay
Hanul Shin, Jung Kwon Lee, Jaehong Kim, and Jiwon Kim · 2017
Earlier work this paper cites.
Continual learning with tiny episodic memories
Arslan Chaudhry, Marcus Rohrbach, Mohamed Elhoseiny, Thalaiyasingam Ajanthan, P Dokania, P Torr, and M Ranzato · 2019
Earlier work this paper cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, Alexander H. Miller, and Sebastian Riedel · 2019
Earlier work this paper cites.
Generalisation guarantees for continual learning with orthogonal gradient descent
Mehdi Abbana Bennani, Thang Doan, and Masashi Sugiyama · 2020
Earlier work this paper cites.
Orthogonal gradient descent for continual learning
Mehrdad Farajtabar, Navid Azizan, Alex Mott, and Ang Li · 2020
Earlier work this paper cites.
Does learning require memorization? A short tale about a long tail
Vitaly Feldman · 2020
Earlier work this paper cites.
The Pile: An 800GB dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy · 2020
Earlier work this paper cites.
Interpreting GPT: the Logit Lens
nostalgebraist · 2020
Earlier work this paper cites.
Distill and replay for continual language learning
Jingyuan Sun, Shaonan Wang, Jiajun Zhang, and Chengqing Zong · 2020
Earlier work this paper cites.
Continual learning with hypernetworks
Johannes von Oswald, Christian Henning, Joao Sacramento, and Benjamin F Grewe · 2020
Earlier work this paper cites.
Modifying memories in transformer models
Chen Zhu, Ankit Singh Rawat, Manzil Zaheer, Srinadh Bhojanapalli, Daliang Li, Felix Yu, and Sanjiv Kumar · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman · 2021
Cited alongside, same era.
A theoretical analysis of catastrophic forgetting through the NTK overlap matrix
Thang Doan, Mehdi Bennani, Bogdan Mazoure, Guillaume Rabusseau, and Pierre Alquier · 2021
Cited alongside, same era.
Amnesiac machine learning
Laura Graves, Vineel Nagisetty, and Vijay Ganesh · 2021
Cited alongside, same era.
Continual learning in the teacher-student setup: Impact of task similarity
Sebastian Lee, Sebastian Goldt, and Andrew Saxe · 2021
Cited alongside, same era.
Anatomy of catastrophic forgetting: Hidden representations and task semantics
Vinay Venkatesh Ramasesh, Ethan Dyer, and Maithra Raghu · 2021
Cited alongside, same era.
Simple entity-centric questions challenge dense retrievers
Continual pre-training of language models
Zixuan Ke, Yijia Shao, Haowei Lin, Tatsuya Konishi, Gyuhak Kim, and Bing Liu · 2023
Later among the works it cites.
When not to trust language models: Investigating effectiveness of parametric and non-parametric memories
Alex Mallen, Akari Asai, Victor Zhong, Rajarshi Das, Daniel Khashabi, and Hannaneh Hajishirzi · 2023
Later among the works it cites.
Mass-editing memory in a transformer
Kevin Meng, Arnab Sen Sharma, Alex J Andonian, Yonatan Belinkov, and David Bau · 2023
Later among the works it cites.
ResMem: Learn what you can and memorize the rest
Zitong Yang, Michal Lukasik, Vaishnavh Nagarajan, Zonglin Li, Ankit Singh Rawat, Manzil Zaheer, Aditya Krishna Menon, and Sanjiv Kumar · 2023
Later among the works it cites.
Counterfactual memorization in neural language models
Chiyuan Zhang, Daphne Ippolito, Katherine Lee, Matthew Jagielski, Florian Tramèr, and Nicholas Carlini · 2023
Later among the works it cites.
MQuAKE: Assessing knowledge editing in language models via multi-hop questions
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Christopher Sciavolino, Zexuan Zhong, Jinhyuk Lee, and Danqi Chen · 2021
Cited alongside, same era.
A Review on Language Models as Knowledge Bases
Badr AlKhamissi, Millicent Li, Asli Celikyilmaz, Mona Diab, and Marjan Ghazvininejad · 2022
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Cited alongside, same era.
How catastrophic can catastrophic forgetting be in linear regression?
Itay Evron, Edward Moroshko, Rachel Ward, Nati Srebro, and Daniel Soudry · 2022
Cited alongside, same era.
LoRA: Low-rank adaptation of large language models
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen · 2022
Cited alongside, same era.
Towards continual knowledge learning of language models
Joel Jang, Seonghyeon Ye, Sohee Yang, Joongbo Shin, Janghoon Han, Gyeonghun Kim, Stanley Jungkyu Choi, and Minjoon Seo · 2022
Cited alongside, same era.
Locating and editing factual associations in GPT
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2022
Cited alongside, same era.
Zexuan Zhong, Zhengxuan Wu, Christopher D Manning, Christopher Potts, and Danqi Chen · 2023
Later among the works it cites.
Physics of language models: Part 3.1, knowledge storage and extraction
Zeyuan Allen-Zhu and Yuanzhi Li · 2024
Closest in time.
Evaluating the ripple effects of knowledge editing in language models
Roi Cohen, Eden Biran, Ori Yoran, Amir Globerson, and Mor Geva · 2024
Closest in time.
The Llama 3 herd of models
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
Closest in time.
Query of CC: Unearthing large scale domain-specific knowledge from public corpora
Zhaoye Fei, Yunfan Shao, Linyang Li, Zhiyuan Zeng, Hang Yan, Xipeng Qiu, and Dahua Lin · 2024
Closest in time.
Does fine-tuning LLMs on new knowledge encourage hallucinations?
Zorik Gekhman, Gal Yona, Roee Aharoni, Matan Eyal, Amir Feder, Roi Reichart, and Jonathan Herzig · 2024
Closest in time.
Understanding catastrophic forgetting in language models via implicit inference
Suhas Kotha, Jacob Mitchell Springer, and Aditi Raghunathan · 2024
Closest in time.
The FineWeb Datasets: Decanting the web for the finest text data at scale
Guilherme Penedo, Hynek Kydlíček, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, Thomas Wolf, et al · 2024
Closest in time.
Train-attention: Meta-learning where to focus in continual knowledge learning
Yeongbin Seo, Dongha Lee, and Jinyoung Yeo · 2024
Closest in time.
Continual learning of large language models: A comprehensive survey
Haizhou Shi, Zihao Xu, Hengyi Wang, Weiyi Qin, Wenyuan Wang, Yibin Wang, and Hao Wang · 2024
Closest in time.
Continual learning for large language models: A survey
Tongtong Wu, Linhao Luo, Yuan-Fang Li, Shirui Pan, Thuy-Trang Vu, and Gholamreza Haffari · 2024
Closest in time.
Investigating continual pretraining in large language models: Insights and implications
Cagatay Yildiz, Nishaanth Kanna Ravichandran, Prishruit Punia, Matthias Bethge, and Beyza Ermis · 2024
Closest in time.
Knowledge overshadowing causes amalgamated hallucination in large language models
Yuji Zhang, Sha Li, Jiateng Liu, Pengfei Yu, Yi R. Fung, Jing Li, Manling Li, and Heng Ji · 2024
Closest in time.
Physics of language models: Part 3.2, knowledge manipulation
Zeyuan Allen-Zhu and Yuanzhi Li · 2025
Closest in time.
Unfamiliar finetuning examples control how language models hallucinate
Katie Kang, Eric Wallace, Claire Tomlin, Aviral Kumar, and Sergey Levine · 2025
Closest in time.
Synthetic continued pretraining
Zitong Yang, Neil Band, Shuangping Li, Emmanuel Candès, and Tatsunori Hashimoto · 2025
Closest in time.