Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are frequently fine-tuned or unlearned to adapt to new tasks or eliminate undesirable behaviors.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, and 1 others. 2020 · 2020
Earlier work this paper cites.
Locating and editing factual associations in gpt
Kevin Meng and 1 others. 2022 · 2022
Earlier work this paper cites.
Emergent abilities of large language models
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, and 1 others. 2022 · 2022
Earlier work this paper cites.
Language models can explain neurons in language models
Steven Bills, Nick Cammarata, Dan Mossing, Henk Tillman, Leo Gao, Gabriel Goh, Ilya Sutskever, Jan Leike, Jeff Wu, and William Saunders. 2023 · 2023
Earlier work this paper cites.
Towards monosemanticity: Decomposing language models with sparse autoencoders
Thomas Bricken, Nelson Conmy, Eric Lieberum, Nelson Elhage, Catherine Olsson, Neel Nanda, Nicholas Joseph, and 1 others. 2023 · 2023
Earlier work this paper cites.
Free dolly: Introducing the world’s first truly open instruction-tuned llm
Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin. 2023 · 2023
Earlier work this paper cites.
Sparse autoencoders find highly interpretable directions in language models
Edward Cunningham, Thibault Sellam, Tal Linzen, and Yonatan Belinkov. 2023 · 2023
Earlier work this paper cites.
Out-of-distribution unlearning: Comprehensive benchmark and analysis
Yue Huang and 1 others. 2023 · 2023
Earlier work this paper cites.
Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks
Samyak Jain, Robert Kirk, Ekdeep Singh Lubana, Robert P Dick, Hidenori Tanaka, Edward Grefenstette, Tim Rocktäschel, and David Scott Krueger. 2023 · 2023
Earlier work this paper cites.
Preserving privacy through dememorization: An unlearning technique for mitigating memorization risks in language models
Aly Kassem, Omar Mahmoud, and Sherif Saad. 2023 · 2023
Earlier work this paper cites.
Unlearning with knowledge distillation in large language models
Minghao Pan and 1 others. 2023 · 2023
Earlier work this paper cites.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. 2023 · 2023
Earlier work this paper cites.
Alpaca: A strong, replicable instruction-following model
Rohan Taori, Ishaan Gulrajani, Tianyi Zhang, Yann Dubois, Xuechen Li, Carlos Guestrin, Percy Liang, and Tatsunori B Hashimoto. 2023 · 2023
Earlier work this paper cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, and 1 others. 2023 · 2023
Earlier work this paper cites.
Forget-me-not: A machine unlearning benchmark for language models
Yujia Xu and 1 others. 2023 · 2023
Cited alongside, same era.
Rejection tuning: Safely unlearning unwanted behaviors in language models
Qiang Yu and 1 others. 2023 · 2023
Cited alongside, same era.
Open source sparse autoencoders for all residual stream layers of gpt2 small
Joseph Bloom. 2024 · 2024
Cited alongside, same era.
Oversampling a topic in the sae training set results in more detailed features related to that topic
Trenton Bricken, Jonathan Marcus, Kelley Rivoire, and Thomas Henighan. 2024a · 2024
Cited alongside, same era.
Stage-wise model diffing
Trenton Bricken, Siddharth Mishra-Sharma, Jonathan Marcus, Adam Jermyn, Christopher Olah, Kelley Rivoire, and Thomas Henighan. 2024b · 2024
Cited alongside, same era.
Automatically interpreting millions of features in large language models
Gonçalo Paulo, Alex Mallen, Caden Juang, and Nora Belrose. 2024 · 2024
Later among the works it cites.
Efficient model editing at scale
Baolin Peng and 1 others. 2024 · 2024
Later among the works it cites.
Fine-tuning enhances existing mechanisms: A case study on entity tracking
Nikhil Prakash, Tamar Rott Shaham, Tal Haklay, Yonatan Belinkov, and David Bau. 2024 · 2024
Later among the works it cites.
Muse: Machine unlearning six-way evaluation for language models
Weijia Shi, Jaechan Lee, Yangsibo Huang, Sadhika Malladi, Jieyu Zhao, Ari Holtzman, Daogao Liu, Luke Zettlemoyer, Noah A Smith, and Chiyuan Zhang. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bart Bussmann, Patrick Leask, and Neel Nanda. 2024 · 2024
Cited alongside, same era.
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, and 1 others. 2024 · 2024
Cited alongside, same era.
Model editing harms general abilities of large language models: Regularization to the rescue
Jia-Chen Gu, Hao-Xiang Xu, Jun-Yu Ma, Pan Lu, Zhen-Hua Ling, Kai-Wei Chang, and Nanyun Peng. 2024 · 2024
Cited alongside, same era.
Dissecting fine-tuning unlearning in large language models
Yihuai Hong, Yuelin Zou, Lijie Hu, Ziqian Zeng, Di Wang, and Haiqin Yang. 2024 · 2024
Cited alongside, same era.
Rwku: Real-world knowledge unlearning in large language models
Zhiwei Jin and 1 others. 2024 · 2024
Cited alongside, same era.
The wmdp benchmark: Measuring and reducing malicious use with unlearning
Nathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue, Daniel Berrios, Alice Gatti, Justin D Li, Ann-Kathrin Dombrowski, Shashwat Goel, Long Phan, and 1 others. 2024 · 2024
Cited alongside, same era.
Sparse crosscoders for cross-layer features and model diffing
Jack Lindsey, Adly Templeton, Jonathan Marcus, Thomas Conerly, Joshua Batson, and Christopher Olah. 2024 · 2024
Cited alongside, same era.
Bozhong Tian, Xiaozhuan Liang, Siyuan Cheng, Qingbin Liu, Mengru Wang, Dianbo Sui, Xi Chen, Huajun Chen, and Ningyu Zhang. 2024 · 2024
Later among the works it cites.
Mmlu-pro: A more robust and challenging multi-task language understanding benchmark
Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, and 1 others. 2024 · 2024
Later among the works it cites.
Unveiling the generalization power of fine-tuned large language models
Haoran Yang, Yumeng Zhang, Jiaqi Xu, Hongyuan Lu, Pheng-Ann Heng, and Wai Lam. 2024 · 2024
Later among the works it cites.
Cross-task generalization abilities of large language models
Qinyuan Ye. 2024 · 2024
Later among the works it cites.
Towards comprehensive post safety alignment of large language models via safety patching
Weixiang Zhao, Yulin Hu, Zhuojun Li, Yang Deng, Jiahe Guo, Xingyu Sui, Yanyan Zhao, Bing Qin, Tat-Seng Chua, and Ting Liu. 2024 · 2024
Later among the works it cites.
Yujie Zhu and 1 others. 2023 · 2024
Later among the works it cites.
Emergent misalignment: Narrow finetuning can produce broadly misaligned llms
Jan Betley, Daniel Tan, Niels Warncke, Anna Sztyber-Betley, Xuchan Bao, Martín Soto, Nathan Labenz, and Owain Evans. 2025 · 2025
Closest in time.
Ai privacy risks & mitigations – large language models (llms)
European Data Protection Board. 2025 · 2025
Closest in time.
Saebench: A comprehensive benchmark for sparse autoencoders in language model interpretability
Adam Karvonen, Can Rager, Johnny Lin, Curt Tigges, Joseph Bloom, David Chanin, Yeu-Tong Lau, Eoin Farrell, Callum McDougall, Kola Ayonrinde, and 1 others. 2025 · 2025
Closest in time.
Robustly identifying concepts introduced during chat fine-tuning using crosscoders
Julian Minder, Clément Dumas, Caden Juang, Bilal Chugtai, and Neel Nanda. 2025 · 2025
Closest in time.