Fetching the paper…
Reading the bibliography…
Our goal is to understand how post-training methods, such as fine-tuning, alignment, and unlearning, modify language model behavior and representations.
Human Error
James Reason · 1990
Earlier work this paper cites.
Measuring massive multitask language understanding, 2021
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2009
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2017
Diederik P. Kingma and Jimmy Ba · 2017
Earlier work this paper cites.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2021
Earlier work this paper cites.
Editing models with task arithmetic
Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi · 2022
Earlier work this paper cites.
Knowledge unlearning for mitigating privacy risks in language models
Joel Jang, Dongkeun Yoon, Sohee Yang, Sungmin Cha, Moontae Lee, Lajanugen Logeswaran, and Minjoon Seo · 2022
Earlier work this paper cites.
Privacy adhering machine un-learning in nlp
Vinayshekhar Bannihatti Kumar, Rashmi Gangadharaiah, and Dan Roth · 2022
Earlier work this paper cites.
Continual learning and private unlearning
Bo Liu, Qiang Liu, and Peter Stone · 2022
Earlier work this paper cites.
Quark: Controllable text generation with reinforced unlearning
Ximing Lu, Sean Welleck, Jack Hessel, Liwei Jiang, Lianhui Qin, Peter West, Prithviraj Ammanabrolu, and Yejin Choi · 2022
Earlier work this paper cites.
Who’s harry potter? approximate unlearning in llms, 2023
Ronen Eldan and Mark Russinovich · 2023
Earlier work this paper cites.
A framework for few-shot language model evaluation, 12 2023
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou · 2023
Earlier work this paper cites.
Knowledge sanitization of large language models
Yoichi Ishibashi and Hidetoshi Shimodaira · 2023
Earlier work this paper cites.
In-context unlearning: Language models as few shot unlearners
Martin Pawelczyk, Seth Neel, and Himabindu Lakkaraju · 2023
Earlier work this paper cites.
Zephyr: Direct distillation of lm alignment, 2023
Lewis Tunstall, Edward Beeching, Nathan Lambert, Nazneen Rajani, Kashif Rasul, Younes Belkada, Shengyi Huang, Leandro von Werra, Clémentine Fourrier, Nathan Habib, Nathan Sarrazin, Omar Sanseviero, Alexander M. Rush, and Thomas Wolf · 2023
Cited alongside, same era.
Large language model unlearning
Yuanshun Yao, Xiaojun Xu, and Yang Liu · 2023
Cited alongside, same era.
Forget-me-not: Learning to forget in text-to-image diffusion models
Eric Zhang, Kai Wang, Xingqian Xu, Zhangyang Wang, and Humphrey Shi · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson · 2023
Cited alongside, same era.
Can sensitive information be deleted from llms? objectives for defending against extraction attacks
Vaidehi Patil, Peter Hase, and Mohit Bansal · 2024
Later among the works it cites.
Representation noising: A defence mechanism against harmful finetuning, 2024
Domenic Rosati, Jan Wehner, Kai Williams, Łukasz Bartoszcze, David Atanasov, Robie Gonzales, Subhabrata Majumdar, Carsten Maple, Hassan Sajjad, and Frank Rudzicz · 2024
Later among the works it cites.
Rethinking llm memorization through the lens of adversarial compression
Avi Schwarzschild, Zhili Feng, Pratyush Maini, Zachary C Lipton, and J Zico Kolter · 2024
Later among the works it cites.
Guardrail baselines for unlearning in llms
Pratiksha Thaker, Yash Maurya, and Virginia Smith · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Andy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka, Nina Panickssery, Wes Gurnee, and Neel Nanda · 2024
Cited alongside, same era.
Model manipulation attacks enable more rigorous evaluations of LLM unlearning
Zora Che, Stephen Casper, Anirudh Satheesh, Rohit Gandikota, Domenic Rosati, Stewart Slocum, Lev E McKinney, Zichu Wu, Zikui Cai, Bilal Chughtai, Furong Huang, and Dylan Hadfield-Menell · 2024
Cited alongside, same era.
Stress-testing capability elicitation with password-locked models
Ryan Greenblatt, Fabien Roger, Dmitrii Krasheninnikov, and David Krueger · 2024
Cited alongside, same era.
Jogging the memory of unlearned model through targeted relearning attack
Shengyuan Hu, Yiwei Fu, Zhiwei Steven Wu, and Virginia Smith · 2024
Cited alongside, same era.
What makes and breaks safety fine-tuning? a mechanistic study, 2024
Samyak Jain, Ekdeep Singh Lubana, Kemal Oksuz, Tom Joy, Philip H. S. Torr, Amartya Sanyal, and Puneet K. Dokania · 2024
Cited alongside, same era.
Soul: Unlocking the power of second-order optimization for llm unlearning
Jinghan Jia, Yihua Zhang, Yimeng Zhang, Jiancheng Liu, Bharat Runwal, James Diffenderfer, Bhavya Kailkhura, and Sijia Liu · 2024
Cited alongside, same era.
Split, unlearn, merge: Leveraging data attributes for more effective unlearning in llms
Swanand Ravindra Kadhe, Farhan Ahmed, Dennis Wei, Nathalie Baracaldo, and Inkit Padhi · 2024
Cited alongside, same era.
Eight methods to evaluate robust unlearning in llms
Aengus Lynch, Phillip Guo, Aidan Ewart, Stephen Casper, and Dylan Hadfield-Menell · 2024
Cited alongside, same era.
Yu Wang, Ruihan Wu, Zexue He, Xiusi Chen, and Julian McAuley · 2024
Later among the works it cites.
What makes unlearning hard and what to do about it, 2024
Kairan Zhao, Meghdad Kurmanji, George-Octavian Bărbulescu, Eleni Triantafillou, and Peter Triantafillou · 2024
Later among the works it cites.
Improving alignment and robustness with circuit breakers, 2024
Andy Zou, Long Phan, Justin Wang, Derek Duenas, Maxwell Lin, Maksym Andriushchenko, Rowan Wang, Zico Kolter, Matt Fredrikson, and Dan Hendrycks · 2024
Later among the works it cites.
Emergent misalignment: Narrow finetuning can produce broadly misaligned llms, 2025
Jan Betley, Daniel Tan, Niels Warncke, Anna Sztyber-Betley, Xuchan Bao, Martín Soto, Nathan Labenz, and Owain Evans · 2025
Closest in time.
Do unlearning methods remove information from language model weights?, 2025
Aghyad Deeb and Fabien Roger · 2025
Closest in time.
Avoiding copyright infringement via large language model unlearning, 2025
Guangyao Dou, Zheyuan Liu, Qing Lyu, Kaize Ding, and Eric Wong · 2025
Closest in time.
Simplicity prevails: Rethinking negative preference optimization for llm unlearning, 2025
Chongyu Fan, Jiancheng Liu, Licong Lin, Jinghan Jia, Ruiqi Zhang, Song Mei, and Sijia Liu · 2025
Closest in time.
Latent adversarial training improves robustness to persistent harmful behaviors in LLMs, 2025
Abhay Sheshadri, Aidan Ewart, Phillip Huang Guo, Aengus Lynch, Cindy Wu, Vivek Hebbar, Henry Sleight, Asa Cooper Stickland, Ethan Perez, Dylan Hadfield-Menell, and Stephen Casper · 2025
Closest in time.
Tamper-resistant safeguards for open-weight LLMs
Rishub Tamirisa, Bhrugu Bharathi, Long Phan, Andy Zhou, Alice Gatti, Tarun Suresh, Maxwell Lin, Justin Wang, Rowan Wang, Ron Arel, Andy Zou, Dawn Song, Bo Li, Dan Hendrycks, and Mantas Mazeika · 2025
Closest in time.