Fetching the paper…
Reading the bibliography…
Methods for knowledge editing and unlearning in large language models seek to edit or remove undesirable knowledge or capabilities without compromising general language modeling performance.
The algorithmic foundations of differential privacy
Cynthia Dwork, Aaron Roth, et al · 2014
Earlier work this paper cites.
Towards making systems forget with machine unlearning
Yinzhi Cao and Junfeng Yang · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization, 2017
Diederik P. Kingma and Jimmy Ba · 2017
Earlier work this paper cites.
Understanding intermediate layers using linear classifier probes, 2018
Guillaume Alain and Yoshua Bengio · 2018
Earlier work this paper cites.
Making ai forget you: Data deletion in machine learning, 2019
Antonio Ginart, Melody Y. Guan, Gregory Valiant, and James Zou · 2019
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling, 2020
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy · 2020
Earlier work this paper cites.
Zoom in: An introduction to circuits
Chris Olah, Nick Cammarata, Ludwig Schubert, Gabriel Goh, Michael Petrov, and Shan Carter · 2020
Earlier work this paper cites.
A mathematical framework for transformer circuits
Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding, 2021
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2021
Earlier work this paper cites.
The power of scale for parameter-efficient prompt tuning, 2021
Brian Lester, Rami Al-Rfou, and Noah Constant · 2021
Earlier work this paper cites.
Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022
Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan · 2022
Earlier work this paper cites.
Discovering knowledge-critical subnetworks in pretrained language models, 2023
Deniz Bayazit, Negar Foroutan, Zeming Chen, Gail Weiss, and Antoine Bosselut · 2023
Earlier work this paper cites.
Pythia: A suite for analyzing large language models across training and scaling, 2023
Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar van der Wal · 2023
Earlier work this paper cites.
Towards automated circuit discovery for mechanistic interpretability, 2023
Arthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch, Stefan Heimersheim, and Adrià Garriga-Alonso · 2023
Earlier work this paper cites.
Dissecting recall of factual associations in auto-regressive language models, 2023
Mor Geva, Jasmijn Bastings, Katja Filippova, and Amir Globerson · 2023
Earlier work this paper cites.
Localizing model behavior with path patching, 2023
Nicholas Goldowsky-Dill, Chris MacLeod, Lucas Sato, and Aryaman Arora · 2023
Cited alongside, same era.
Does localization inform editing? surprising differences in causality-based localization vs. knowledge editing in language models, 2023
Peter Hase, Mohit Bansal, Been Kim, and Asma Ghandeharioun · 2023
Cited alongside, same era.
Self-destructing models: Increasing the costs of harmful dual uses of foundation models
Peter Henderson, Eric Mitchell, Christopher Manning, Dan Jurafsky, and Chelsea Finn · 2023
Cited alongside, same era.
Surgical fine-tuning improves adaptation to distribution shifts, 2023
Yoonho Lee, Annie S. Chen, Fahim Tajwar, Ananya Kumar, Huaxiu Yao, Percy Liang, and Chelsea Finn · 2023
Cited alongside, same era.
Lora fine-tuning efficiently undoes safety training in llama 2-chat 70b, 2023
Simon Lermen, Charlie Rogers-Smith, and Jeffrey Ladish · 2023
Cited alongside, same era.
Summing up the facts: Additive mechanisms behind factual recall in llms, 2024
Bilal Chughtai, Alan Cooney, and Neel Nanda · 2024
Closest in time.
Do unlearning methods remove information from language model weights?, 2024
Aghyad Deeb and Fabien Roger · 2024
Closest in time.
Sophon: Non-fine-tunable learning to restrain task transferability for pre-trained models
Jiangyi Deng, Shengyuan Pang, Yanjiao Chen, Liangming Xia, Yijie Bai, Haiqin Weng, and Wenyuan Xu · 2024
Closest in time.
Intrinsic evaluation of unlearning using parametric knowledge traces, 2024
Yihuai Hong, Lei Yu, Haiqin Yang, Shauli Ravfogel, and Mor Geva · 2024
Closest in time.
A mechanistic understanding of alignment algorithms: A case study on dpo and toxicity
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Locating and editing factual associations in gpt, 2023
Kevin Meng, David Bau, Alex Andonian, and Yonatan Belinkov · 2023
Cited alongside, same era.
Attribution patching: Activation patching at industrial scale, 2023
Neel Nanda · 2023
Cited alongside, same era.
Fact finding: Attempting to reverse-engineer factual recall on the neuron level, Dec 2023
Neel Nanda, Senthooran Rajamanoharan, János Kramár, and Rohin Shah · 2023
Cited alongside, same era.
Task-specific skill localization in fine-tuned language models, 2023
Abhishek Panigrahi, Nikunj Saunshi, Haoyu Zhao, and Sanjeev Arora · 2023
Cited alongside, same era.
Can sensitive information be deleted from llms? objectives for defending against extraction attacks, 2023
Vaidehi Patil, Peter Hase, and Mohit Bansal · 2023
Cited alongside, same era.
Attribution patching outperforms automated circuit discovery, 2023
Aaquib Syed, Can Rager, and Arthur Conmy · 2023
Cited alongside, same era.
Machine unlearning: A survey
Heng Xu, Tianqing Zhu, Lefeng Zhang, Wanlei Zhou, and Philip S Yu · 2023
Cited alongside, same era.
Andrew Lee, Xiaoyan Bai, Itamar Pres, Martin Wattenberg, Jonathan K Kummerfeld, and Rada Mihalcea · 2024
Closest in time.
Large language models relearn removed concepts, 2024
Michelle Lo, Shay B. Cohen, and Fazl Barez · 2024
Closest in time.
Eight methods to evaluate robust unlearning in llms, 2024
Aengus Lynch, Phillip Guo, Aidan Ewart, Stephen Casper, and Dylan Hadfield-Menell · 2024
Closest in time.
Harmbench: A standardized evaluation framework for automated red teaming and robust refusal, 2024
Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, David Forsyth, and Dan Hendrycks · 2024
Closest in time.
tinybenchmarks: evaluating llms with fewer examples, 2024
Felipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun, Gongjun Xu, and Mikhail Yurochkin · 2024
Closest in time.
Tamper-resistant safeguards for open-weight llms, 2024
Rishub Tamirisa, Bhrugu Bharathi, Long Phan, Andy Zhou, Alice Gatti, Tarun Suresh, Maxwell Lin, Justin Wang, Rowan Wang, Ron Arel, Andy Zou, Dawn Song, Bo Li, Dan Hendrycks, and Mantas Mazeika · 2024
Closest in time.
Scaling monosemanticity: Extracting interpretable features from claude 3 sonnet
Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L Turner, Callum McDougall, Monte MacDiarmid, C. Daniel Freeman, Theodore R. Sumers, Edward Rees, Joshua Batson, Adam Jermyn, Shan Carter, Chris Olah, and Tom Henighan · 2024
Closest in time.
In-context learning can re-learn forbidden tasks
Sophie Xhonneux, David Dobre, Jian Tang, Gauthier Gidel, and Dhanya Sridhar · 2024
Closest in time.
Low-resource languages jailbreak gpt-4, 2024
Zheng-Xin Yong, Cristina Menghini, and Stephen H. Bach · 2024
Closest in time.
Improving alignment and robustness with circuit breakers, 2024
Andy Zou, Long Phan, Justin Wang, Derek Duenas, Maxwell Lin, Maksym Andriushchenko, Rowan Wang, Zico Kolter, Matt Fredrikson, and Dan Hendrycks · 2024
Closest in time.