Fetching the paper…
Reading the bibliography…
Machine unlearning can be useful for removing harmful capabilities and memorized text from large language models (LLMs), but there are not yet standardized methods for rigorously evaluating it.
Harry potter series
J.K. Rowling · 2007
Earlier work this paper cites.
Towards making systems forget with machine unlearning
Yinzhi Cao and Junfeng Yang · 2015
Earlier work this paper cites.
Certified data removal from machine learning models
Chuan Guo, Tom Goldstein, Awni Hannun, and Laurens Van Der Maaten · 2019
Earlier work this paper cites.
Machine unlearning
Lucas Bourtoule, Varun Chandrasekaran, Christopher A Choquette-Choo, Hengrui Jia, Adelin Travers, Baiwu Zhang, David Lie, and Nicolas Papernot · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
J. Edward Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, and Weizhu Chen · 2021
Earlier work this paper cites.
Effect of scale on catastrophic forgetting in neural networks
Vinay Venkatesh Ramasesh, Aitor Lewkowycz, and Ethan Dyer · 2021
Earlier work this paper cites.
Probing classifiers: Promises, shortcomings, and advances
Yonatan Belinkov · 2022
Earlier work this paper cites.
Discovering latent knowledge in language models without supervision
Collin Burns, Haotian Ye, Dan Klein, and Jacob Steinhardt · 2022
Earlier work this paper cites.
Quantifying memorization across neural language models
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang · 2022
Earlier work this paper cites.
Continual pre-training mitigates forgetting in language and vision
Andrea Cossu, Tinne Tuytelaars, Antonio Carta, Lucia Passaro, Vincenzo Lomonaco, and Davide Bacciu · 2022
Earlier work this paper cites.
Editing models with task arithmetic
Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Suchin Gururangan, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi · 2022
Earlier work this paper cites.
Knowledge unlearning for mitigating privacy risks in language models
Joel Jang, Dongkeun Yoon, Sohee Yang, Sungmin Cha, Moontae Lee, Lajanugen Logeswaran, and Minjoon Seo · 2022
Earlier work this paper cites.
Linear connectivity reveals generalization strategies
Jeevesh Juneja, Rachit Bansal, Kyunghyun Cho, João Sedoc, and Naomi Saphra · 2022
Earlier work this paper cites.
Privacy adhering machine un-learning in nlp
Vinayshekhar Bannihatti Kumar, Rashmi Gangadharaiah, and Dan Roth · 2022
Earlier work this paper cites.
Technical report for iccv 2021 challenge sslad-track3b: Transformers are better continual learners
Duo Li, Guimei Cao, Yunlu Xu, Zhanzhan Cheng, and Yi Niu · 2022
Earlier work this paper cites.
Quark: Controllable text generation with reinforced unlearning
Ximing Lu, Sean Welleck, Jack Hessel, Liwei Jiang, Lianhui Qin, Peter West, Prithviraj Ammanabrolu, and Yejin Choi · 2022
Earlier work this paper cites.
A survey of machine unlearning
Thanh Tam Nguyen, Thanh Trung Huynh, Phi Le Nguyen, Alan Wee-Chung Liew, Hongzhi Yin, and Quoc Viet Hung Nguyen · 2022
Earlier work this paper cites.
Fine-tuned language models are continual learners
Thomas Scialom, Tuhin Chakrabarty, and Smaranda Muresan · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
Unlearn what you want to forget: Efficient unlearning for llms
Jiaao Chen and Diyi Yang · 2023
Earlier work this paper cites.
Who’s harry potter? approximate unlearning in llms
Ronen Eldan and Mark Russinovich · 2023
Earlier work this paper cites.
Language models represent space and time
Wes Gurnee and Max Tegmark · 2023
Cited alongside, same era.
Self-destructing models: Increasing the costs of harmful dual uses of foundation models
Peter Henderson, Eric Mitchell, Christopher Manning, Dan Jurafsky, and Chelsea Finn · 2023
Cited alongside, same era.
Knowledge sanitization of large language models
Yoichi Ishibashi and Hidetoshi Shimodaira · 2023
Cited alongside, same era.
Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks
Samyak Jain, Robert Kirk, Ekdeep Singh Lubana, Robert P Dick, Hidenori Tanaka, Edward Grefenstette, Tim Rocktäschel, and David Scott Krueger · 2023
Cited alongside, same era.
Activation addition: Steering language models without optimization
Alex Turner, Lisa Thiergart, David Udell, Gavin Leech, Ulisse Mini, and Monte MacDiarmid · 2023
Later among the works it cites.
Kga: A general machine unlearning framework based on knowledge gap alignment
Lingzhi Wang, Tong Chen, Wei Yuan, Xingshan Zeng, Kam-Fai Wong, and Hongzhi Yin · 2023
Later among the works it cites.
Jailbroken: How does llm safety training fail?
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
Later among the works it cites.
Depn: Detecting and editing privacy neurons in pretrained language models
Xinwei Wu, Junzhuo Li, Minghui Xu, Weilong Dong, Shuangzhi Wu, Chao Bian, and Deyi Xiong · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Antonia Karamolegkou, Jiaang Li, Li Zhou, and Anders Søgaard · 2023
Cited alongside, same era.
Understanding catastrophic forgetting in language models via implicit inference
Suhas Kotha, Jacob Mitchell Springer, and Aditi Raghunathan · 2023
Cited alongside, same era.
Lora fine-tuning efficiently undoes safety training in llama 2-chat 70b
Simon Lermen, Charlie Rogers-Smith, and Jeffrey Ladish · 2023
Cited alongside, same era.
Mechanistic mode connectivity
Ekdeep Singh Lubana, Eric J Bigelow, Robert P Dick, David Krueger, and Hidenori Tanaka · 2023
Cited alongside, same era.
Investigating forgetting in pre-trained representations through continual learning
Yun Luo, Zhen Yang, Xuefeng Bai, Fandong Meng, Jie Zhou, and Yue Zhang · 2023
Cited alongside, same era.
Can sensitive information be deleted from llms? objectives for defending against extraction attacks
Vaidehi Patil, Peter Hase, and Mohit Bansal · 2023
Cited alongside, same era.
In-context unlearning: Language models as few shot unlearners
Martin Pawelczyk, Seth Neel, and Himabindu Lakkaraju · 2023
Cited alongside, same era.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson · 2023
Cited alongside, same era.
Xianjun Yang, Xiao Wang, Qi Zhang, Linda Petzold, William Yang Wang, Xun Zhao, and Dahua Lin · 2023
Later among the works it cites.
Low-resource languages jailbreak gpt-4
Zheng-Xin Yong, Cristina Menghini, and Stephen H Bach · 2023
Later among the works it cites.
Unlearning bias in language models by partitioning gradients
Charles Yu, Sullam Jeoung, Anish Kasi, Pengfei Yu, and Heng Ji · 2023
Later among the works it cites.
Removing rlhf protections in gpt-4 via fine-tuning
Qiusi Zhan, Richard Fang, Rohan Bindu, Akul Gupta, Tatsunori Hashimoto, and Daniel Kang · 2023
Later among the works it cites.
Composing parameter-efficient modules with arithmetic operations
Jinghan Zhang, Shiqi Chen, Junteng Liu, and Junxian He · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Andy Zou, Zifan Wang, J Zico Kolter, and Matt Fredrikson · 2023
Later among the works it cites.
Coercing llms to do and reveal (almost) anything, 2024
Jonas Geiping, Alex Stein, Manli Shu, Khalid Saifullah, Yuxin Wen, and Tom Goldstein · 2024
Closest in time.
Corrective machine unlearning, 2024
Shashwat Goel, Ameya Prabhu, Philip Torr, Ponnurangam Kumaraguru, and Amartya Sanyal · 2024
Closest in time.
Sleeper agents: Training deceptive llms that persist through safety training
Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M Ziegler, Tim Maxwell, Newton Cheng, et al · 2024
Closest in time.
A mechanistic understanding of alignment algorithms: A case study on dpo and toxicity
Andrew Lee, Xiaoyan Bai, Itamar Pres, Martin Wattenberg, Jonathan K Kummerfeld, and Rada Mihalcea · 2024
Closest in time.
Large language models relearn removed concepts
Michelle Lo, Shay B Cohen, and Fazl Barez · 2024
Closest in time.
Investigating bias representations in llama 2 chat via activation steering, 2024
Dawn Lu and Nina Rimsky · 2024
Closest in time.
Tofu: A task of fictitious unlearning for llms
Pratyush Maini, Zhili Feng, Avi Schwarzschild, Zachary C Lipton, and J Zico Kolter · 2024
Closest in time.
Fine-tuning enhances existing mechanisms: A case study on entity tracking, 2024
Nikhil Prakash, Tamar Rott Shaham, Tal Haklay, Yonatan Belinkov, and David Bau · 2024
Closest in time.
Soft prompt threats: Attacking safety alignment and unlearning in open-source llms through the embedding space, 2024
Leo Schwinn, David Dobre, Sophie Xhonneux, Gauthier Gidel, and Stephan Gunnemann · 2024
Closest in time.
A language model’s guide through latent space, 2024
Dimitri von Rütte, Sotiris Anagnostidis, Gregor Bachmann, and Thomas Hofmann · 2024
Closest in time.
Assessing the brittleness of safety alignment via pruning and low-rank modifications, 2024
Boyi Wei, Kaixuan Huang, Yangsibo Huang, Tinghao Xie, Xiangyu Qi, Mengzhou Xia, Prateek Mittal, Mengdi Wang, and Peter Henderson · 2024
Closest in time.