Fetching the paper…
Reading the bibliography…
Adapting LLMs with new knowledge is increasingly important, but standard fine-tuning often erodes aligned epistemic abstention: the ability to acknowledge when the model does not know.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J Cohen · 1989
Earlier work this paper cites.
Connectionist models of recognition memory: constraints imposed by learning and forgetting functions
Roger Ratcliff · 1990
Earlier work this paper cites.
Catastrophic forgetting, rehearsal and pseudorehearsal
Anthony Robins · 1995
Earlier work this paper cites.
Overcoming catastrophic forgetting in neural networks
James Kirkpatrick, Razvan Pascanu, Neil Rabinowitz, Joel Veness, Guillaume Desjardins, Andrei A Rusu, Kieran Milan, John Quan, Tiago Ramalho, Agnieszka Grabska-Barwinska, et al · 2017
Earlier work this paper cites.
Gradient episodic memory for continual learning
David Lopez-Paz and Marc’Aurelio Ranzato · 2017
Earlier work this paper cites.
Overcoming catastrophic forgetting with hard attention to the task
Joan Serra, Didac Suris, Marius Miron, and Alexandros Karatzoglou · 2018
Earlier work this paper cites.
Continual learning with node-importance based adaptive group sparse regularization
Sangwon Jung, Hongjoon Ahn, Sungmin Cha, and Taesup Moon · 2020
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models. arxiv 2021
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Earlier work this paper cites.
Editing models with task arithmetic
Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi · 2022
Earlier work this paper cites.
An empirical study of catastrophic forgetting in large language models during continual fine-tuning
Yun Luo, Zhen Yang, Fandong Meng, Yafu Li, Jie Zhou, and Yue Zhang · 2023
Earlier work this paper cites.
The linear representation hypothesis and the geometry of large language models
Kiho Park, Yo Joong Choe, and Victor Veitch · 2023
Earlier work this paper cites.
A closer look at rehearsal-free continual learning
James Seale Smith, Junjiao Tian, Shaunak Halbe, Yen-Chang Hsu, and Zsolt Kira · 2023
Cited alongside, same era.
Steering language models with activation engineering
Alexander Matt Turner, Lisa Thiergart, Gavin Leech, David Udell, Juan J Vazquez, Ulisse Mini, and Monte MacDiarmid · 2023
Cited alongside, same era.
Representation engineering: A top-down approach to ai transparency
Andy Zou, Long Phan, Sarah Chen, James Campbell, Phillip Guo, Richard Ren, Alexander Pan, Xuwang Yin, Mantas Mazeika, Ann-Kathrin Dombrowski, et al · 2023
Cited alongside, same era.
I don’t know: Explicit modeling of uncertainty with an [idk] token
Roi Cohen, Konstantin Dobler, Eden Biran, and Gerard de Melo · 2024
Cited alongside, same era.
Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Amy Yang, Angela Fan, et al · 2024
To believe or not to believe your llm
Yasin Abbasi Yadkori, Ilja Kuzborskij, András György, and Csaba Szepesvári · 2024
Later among the works it cites.
R-tuning: Instructing large language models to say ‘i don’t know’
Hanning Zhang, Shizhe Diao, Yong Lin, Yi Fung, Qing Lian, Xingyao Wang, Yangyi Chen, Heng Ji, and Tong Zhang · 2024
Later among the works it cites.
Revolutionizing finance with llms: An overview of applications and insights
Huaqin Zhao, Zhengliang Liu, Zihao Wu, Yiwei Li, Tianze Yang, Peng Shu, Shaochen Xu, Haixing Dai, Lin Zhao, Gengchen Mai, et al · 2024
Later among the works it cites.
Steering out-of-distribution generalization with concept ablation fine-tuning
Helena Casademunt, Caden Juang, Adam Karvonen, Samuel Marks, Senthooran Rajamanoharan, and Neel Nanda · 2025
Closest in time.
Persona vectors: Monitoring and controlling character traits in language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Does fine-tuning llms on new knowledge encourage hallucinations?
Zorik Gekhman, Gal Yona, Roee Aharoni, Matan Eyal, Amir Feder, Roi Reichart, and Jonathan Herzig · 2024
Cited alongside, same era.
Large language models in law: A survey
Jinqi Lai, Wensheng Gan, Jiayang Wu, Zhenlian Qi, and Philip S Yu · 2024
Cited alongside, same era.
Lei Liu, Xiaoyan Yang, Junchi Lei, Xiaoyang Liu, Yue Shen, Zhiqiang Zhang, Peng Wei, Jinjie Gu, Zhixuan Chu, Zhan Qin, et al · 2024
Cited alongside, same era.
Tofu: A task of fictitious unlearning for llms
Pratyush Maini, Zhili Feng, Avi Schwarzschild, Zachary C Lipton, and J Zico Kolter · 2024
Cited alongside, same era.
Pistol: Dataset compilation pipeline for structural unlearning of llms
Xinchi Qiu, William F Shen, Yihong Chen, Nicola Cancedda, Pontus Stenetorp, and Nicholas D Lane · 2024
Cited alongside, same era.
An Yang, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chengyuan Li, Dayiheng Liu, Fei Huang, Haoran Wei, et al
Cited in the paper.
Alignment for honesty
Yuqing Yang, Ethan Chern, Xipeng Qiu, Graham Neubig, and Pengfei Liu
Cited in the paper.
Runjin Chen, Andy Arditi, Henry Sleight, Owain Evans, and Jack Lindsey · 2025
Closest in time.
Calibrating verbal uncertainty as a linear feature to reduce hallucinations
Ziwei Ji, Lei Yu, Yeskendir Koishekenov, Yejin Bang, Anthony Hartshorn, Alan Schelten, Cheng Zhang, Pascale Fung, and Nicola Cancedda · 2025
Closest in time.
Overcoming catastrophic forgetting in neural networks
Brandon Shuen Yi Loke, Filippo Quadri, Gabriel Vivanco, Maximilian Casagrande, and Saúl Fenollosa · 2025
Closest in time.
Controlled low-rank adaptation with subspace regularization for continued training on large language models
Yuheng Lu, Bingshuo Qian, Caixia Yuan, Huixing Jiang, and Xiaojie Wang · 2025
Closest in time.
Lunar: Llm unlearning via neural activation redirection
William F Shen, Xinchi Qiu, Meghdad Kurmanji, Alex Iacob, Lorenzo Sani, Yihong Chen, Nicola Cancedda, and Nicholas D Lane · 2025
Closest in time.
Why representation engineering works: A theoretical and empirical study in vision-language models
Bowei Tian, Xuntao Lyu, Meng Liu, Hongyi Wang, and Ang Li · 2025
Closest in time.