Fetching the paper…
Reading the bibliography…
Current LLM unlearning methods are not robust.
On information and sufficiency
Solomon Kullback and Richard A Leibler · 1951
Earlier work this paper cites.
The right to be forgotten
Jeffrey Rosen · 2011
Earlier work this paper cites.
Towards making systems forget with machine unlearning
Yinzhi Cao and Junfeng Yang · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher · 2016
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang · 2017
Earlier work this paper cites.
Membership inference attacks against machine learning models
Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov · 2017
Earlier work this paper cites.
Making ai forget you: Data deletion in machine learning
Antonio Ginart, Melody Guan, Gregory Valiant, and James Y Zou · 2019
Earlier work this paper cites.
Certified data removal from machine learning models
Chuan Guo, Tom Goldstein, Awni Hannun, and Laurens Van Der Maaten · 2019
Earlier work this paper cites.
On warm-starting neural network training
Jordan Ash and Ryan P Adams · 2020
Earlier work this paper cites.
Formalizing data deletion in the context of the right to be forgotten
Sanjam Garg, Shafi Goldwasser, and Prashant Nalini Vasudevan · 2020
Earlier work this paper cites.
Machine unlearning
Lucas Bourtoule, Varun Chandrasekaran, Christopher A Choquette-Choo, Hengrui Jia, Adelin Travers, Baiwu Zhang, David Lie, and Nicolas Papernot · 2021
Earlier work this paper cites.
Descent-to-delete: Gradient-based methods for machine unlearning
Seth Neel, Aaron Roth, and Saeed Sharifi-Malvajerdi · 2021
Earlier work this paper cites.
Remember what you want to forget: Algorithms for machine unlearning
Ayush Sekhari, Jayadev Acharya, Gautam Kamath, and Ananda Theertha Suresh · 2021
Earlier work this paper cites.
Machine unlearning via algorithmic stability
Enayat Ullah, Tung Mai, Anup Rao, Ryan A Rossi, and Raman Arora · 2021
Earlier work this paper cites.
If influence functions are the answer, then what is the question?
Juhan Bae, Nathan Ng, Alston Lo, Marzyeh Ghassemi, and Roger B Grosse · 2022
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, et al · 2022
Earlier work this paper cites.
A survey of machine unlearning
Thanh Tam Nguyen, Thanh Trung Huynh, Zhao Ren, Phi Le Nguyen, Alan Wee-Chung Liew, Hongzhi Yin, and Quoc Viet Hung Nguyen · 2022
Earlier work this paper cites.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al · 2022
Earlier work this paper cites.
Algorithms that approximate data removal: New results and limitations
Vinith Suriyakumar and Ashia C Wilson · 2022
Earlier work this paper cites.
On the necessity of auditable algorithmic definitions for machine unlearning
Anvith Thudi, Hengrui Jia, Ilia Shumailov, and Nicolas Papernot · 2022
Earlier work this paper cites.
Unlearn what you want to forget: Efficient unlearning for llms
Jiaao Chen and Diyi Yang · 2023
Earlier work this paper cites.
Forget unlearning: Towards true data-deletion in machine learning
Rishav Chourasia and Neil Shah · 2023
Earlier work this paper cites.
Can bad teaching induce forgetting? unlearning in deep networks using an incompetent teacher
Vikram S Chundawat, Ayush K Tarun, Murari Mandal, and Mohan Kankanhalli · 2023
Earlier work this paper cites.
Salun: Empowering machine unlearning via gradient-based weight saliency in both image classification and generation
Chongyu Fan, Jiancheng Liu, Yihua Zhang, Eric Wong, Dennis Wei, and Sijia Liu · 2023
Earlier work this paper cites.
Knowledge unlearning for mitigating privacy risks in language models
Joel Jang, Dongkeun Yoon, Sohee Yang, Sungmin Cha, Moontae Lee, Lajanugen Logeswaran, and Minjoon Seo · 2023
Earlier work this paper cites.
Model sparsity can simplify machine unlearning
Jinghan Jia, Jiancheng Liu, Parikshit Ram, Yuguang Yao, Gaowen Liu, Yang Liu, Pranay Sharma, and Sijia Liu · 2023
Earlier work this paper cites.
Towards unbounded machine unlearning
Meghdad Kurmanji, Peter Triantafillou, Jamie Hayes, and Eleni Triantafillou · 2023
Earlier work this paper cites.
Selective and collaborative influence function for efficient recommendation unlearning
Yuyuan Li, Chaochao Chen, Xiaolin Zheng, Yizhao Zhang, Biao Gong, Jun Wang, and Linxun Chen · 2023
Earlier work this paper cites.
Preparedness framework (beta)
OpenAI · 2023
Cited alongside, same era.
Jonas B Sandbrink · 2023
Cited alongside, same era.
Are emergent abilities of large language models a mirage?
Rylan Schaeffer, Brando Miranda, and Sanmi Koyejo · 2023
Cited alongside, same era.
Knowledge unlearning for llms: Tasks, methods, and challenges
Nianwen Si, Hao Zhang, Heyu Chang, Wenlin Zhang, Dan Qu, and Weiqiang Zhang · 2023
Cited alongside, same era.
Kga: A general machine unlearning framework based on knowledge gap alignment
Lingzhi Wang, Tong Chen, Wei Yuan, Xingshan Zeng, Kam-Fai Wong, and Hongzhi Yin · 2023
Cited alongside, same era.
Can sensitive information be deleted from llms? objectives for defending against extraction attacks
Vaidehi Patil, Peter Hase, and Mohit Bansal · 2024
Later among the works it cites.
Representation noising: A defence mechanism against harmful finetuning
Domenic Rosati, Jan Wehner, Kai Williams, Lukasz Bartoszcze, Robie Gonzales, Subhabrata Majumdar, Hassan Sajjad, Frank Rudzicz, et al · 2024
Later among the works it cites.
Extracting unlearned information from llms with activation steering
Atakan Seyitoğlu, Aleksei Kuvshinov, Leo Schwinn, and Stephan Günnemann · 2024
Later among the works it cites.
Latent adversarial training improves robustness to persistent harmful behaviors in llms
Abhay Sheshadri, Aidan Ewart, Phillip Guo, Aengus Lynch, Cindy Wu, Vivek Hebbar, Henry Sleight, Asa Cooper Stickland, Ethan Perez, Dylan Hadfield-Menell, et al · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
Cited alongside, same era.
Shadow alignment: The ease of subverting safely-aligned language models
Xianjun Yang, Xiao Wang, Qi Zhang, Linda Petzold, William Yang Wang, Xun Zhao, and Dahua Lin · 2023
Cited alongside, same era.
Large language model unlearning
Yuanshun Yao, Xiaojun Xu, and Yang Liu · 2023
Cited alongside, same era.
Foundational challenges in assuring alignment and safety of large language models
Usman Anwar, Abulhair Saparov, Javier Rando, Daniel Paleka, Miles Turpin, Peter Hase, Ekdeep Singh Lubana, Erik Jenner, Stephen Casper, Oliver Sourbut, et al · 2024
Cited alongside, same era.
To each (textual sequence) its own: improving memorized-data unlearning in large language models
George-Octavian Bărbulescu and Peter Triantafillou · 2024
Cited alongside, same era.
Stable lm 2 1.6 b technical report
Marco Bellagente, Jonathan Tow, Dakota Mahan, Duy Phung, Maksym Zhuravinskyi, Reshinth Adithyan, James Baicoianu, Ben Brooks, Nathan Cooper, Ashish Datta, et al · 2024
Cited alongside, same era.
Data poisoning in llms: Jailbreak-tuning and scaling laws
Dillon Bowen, Brendan Murphy, Will Cai, David Khachaturov, Adam Gleave, and Kellin Pelrine · 2024
Cited alongside, same era.
Rishub Tamirisa, Bhrugu Bharathi, Long Phan, Andy Zhou, Alice Gatti, Tarun Suresh, Maxwell Lin, Justin Wang, Rowan Wang, Ron Arel, et al · 2024
Later among the works it cites.
Calvin Tan and Jerome Wang · 2024
Later among the works it cites.
Gemma 2 2b, 2024
Gemma Team · 2024
Later among the works it cites.
Gemma 2: Improving open language models at a practical size
Gemma Team, Morgane Riviere, Shreya Pathak, Pier Giuseppe Sessa, Cassidy Hardin, Surya Bhupatiraju, Léonard Hussenot, Thomas Mesnard, Bobak Shahriari, Alexandre Ramé, et al · 2024
Later among the works it cites.
Position: Llm unlearning benchmarks are weak measures of progress
Pratiksha Thaker, Shengyuan Hu, Neil Kale, Yash Maurya, Zhiwei Steven Wu, and Virginia Smith · 2024
Later among the works it cites.
Towards reliable empirical machine unlearning evaluation: A game-theoretic view
Yiwen Tu, Pingbang Hu, and Jiaqi Ma · 2024
Later among the works it cites.
Bichen Wang, Yuzhe Zi, Yixin Sun, Yanyan Zhao, and Bing Qin · 2024
Later among the works it cites.
Evaluating copyright takedown methods for language models
Boyi Wei, Weijia Shi, Yangsibo Huang, Noah A Smith, Chiyuan Zhang, Luke Zettlemoyer, Kai Li, and Peter Henderson · 2024
Later among the works it cites.
Survey on knowledge distillation for large language models: methods, evaluation, and application
Chuanpeng Yang, Yao Zhu, Wang Lu, Yidong Wang, Qian Chen, Chenlong Gao, Bingjie Yan, and Yiqiang Chen · 2024
Later among the works it cites.
Machine unlearning of pre-trained large language models
Jin Yao, Eli Chien, Minxin Du, Xinyao Niu, Tianhao Wang, Zezhou Cheng, and Xiang Yue · 2024
Later among the works it cites.
What makes unlearning hard and what to do about it
Kairan Zhao, Meghdad Kurmanji, George-Octavian Bărbulescu, Eleni Triantafillou, and Peter Triantafillou · 2024
Later among the works it cites.
Open problems in machine unlearning for ai safety
Fazl Barez, Tingchen Fu, Ameya Prabhu, Stephen Casper, Amartya Sanyal, Adel Bibi, Aidan O’Gara, Robert Kirk, Ben Bucknall, Tim Fist, et al · 2025
Closest in time.
Model tampering attacks enable more rigorous evaluations of llm capabilities
Zora Che, Stephen Casper, Robert Kirk, Anirudh Satheesh, Stewart Slocum, Lev E McKinney, Rohit Gandikota, Aidan Ewart, Domenic Rosati, Zichu Wu, et al · 2025
Closest in time.
Group-robust machine unlearning
Thomas De Min, Subhankar Roy, Stéphane Lathuilière, Elisa Ricci, and Massimiliano Mancini · 2025
Closest in time.
Openunlearning: Accelerating llm unlearning via unified benchmarking of methods and metrics
Vineeth Dorna, Anmol Mekala, Wenlong Zhao, Andrew McCallum, Zachary C Lipton, J Zico Kolter, and Pratyush Maini · 2025
Closest in time.
Chongyu Fan, Jinghan Jia, Yihua Zhang, Anil Ramakrishna, Mingyi Hong, and Sijia Liu · 2025
Closest in time.
Machine unlearning via simulated oracle matching
Kristian Georgiev, Roy Rinberg, Sung Min Park, Shivam Garg, Andrew Ilyas, Aleksander Madry, and Seth Neel · 2025
Closest in time.
The elicitation game: Evaluating capability elicitation techniques
Felix Hofstätter, Teun van der Weij, Jayden Teoh, Henning Bartsch, and Francis Rhys Ward · 2025
Closest in time.
Unlearning or obfuscating? jogging the memory of unlearned LLMs via benign relearning
Shengyuan Hu, Yiwei Fu, Steven Wu, and Virginia Smith · 2025
Closest in time.
Detecting and filtering unsafe training data via data attribution
Yijun Pan, Taiwei Shi, Jieyu Zhao, and Jiaqi W Ma · 2025
Closest in time.
Layered unlearning for adversarial relearning
Timothy Qian, Vinith Suriyakumar, Ashia Wilson, and Dylan Hadfield-Menell · 2025
Closest in time.
Gemma Team, Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre Ramé, Morgane Rivière, et al · 2025
Closest in time.
Erasing without remembering: Safeguarding knowledge forgetting in large language models
Huazheng Wang, Yongcheng Jing, Haifeng Sun, Yingjie Wang, Jingyu Wang, Jianxin Liao, and Dacheng Tao · 2025
Closest in time.
Unlearning isn’t deletion: Investigating reversibility of machine unlearning in llms
Xiaoyu Xu, Xiang Yue, Yang Liu, Qingqing Ye, Haibo Hu, and Minxin Du · 2025
Closest in time.