Fetching the paper…
Reading the bibliography…
Recent work has demonstrated that finetuning is a promising approach to 'unlearn' concepts from large language models.
Towards making systems forget with machine unlearning
Yinzhi Cao and Junfeng Yang · 2015
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal · 2018
Earlier work this paper cites.
Making ai forget you: Data deletion in machine learning
Antonio Ginart, Melody Guan, Gregory Valiant, and James Y Zou · 2019
Earlier work this paper cites.
Making machine learning forget
Saurabh Shintre, Kevin A Roundy, and Jasjeet Dhaliwal · 2019
Earlier work this paper cites.
Certified data removal from machine learning models
Chuan Guo, Tom Goldstein, Awni Hannun, and Laurens Van Der Maaten · 2019
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Earlier work this paper cites.
Vikas Yadav, Steven Bethard, and Mihai Surdeanu · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Earlier work this paper cites.
Eternal sunshine of the spotless net: Selective forgetting in deep networks
Aditya Golatkar, Alessandro Achille, and Stefano Soatto · 2020
Earlier work this paper cites.
{ \{ Updates-Leak } \} : Data set inference and reconstruction attacks in online learning
Ahmed Salem, Apratim Bhattacharya, Michael Backes, Mario Fritz, and Yang Zhang · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2020
Earlier work this paper cites.
Machine unlearning
Lucas Bourtoule, Varun Chandrasekaran, Christopher A Choquette-Choo, Hengrui Jia, Adelin Travers, Baiwu Zhang, David Lie, and Nicolas Papernot · 2021
Earlier work this paper cites.
Adaptive machine unlearning
Varun Gupta, Christopher Jung, Seth Neel, Aaron Roth, Saeed Sharifi-Malvajerdi, and Chris Waites · 2021
Earlier work this paper cites.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le · 2021
Earlier work this paper cites.
When machine unlearning jeopardizes privacy
Min Chen, Zhikun Zhang, Tianhao Wang, Michael Backes, Mathias Humbert, and Yang Zhang · 2021
Earlier work this paper cites.
What does it mean for a language model to preserve privacy?
Hannah Brown, Katherine Lee, Fatemehsadat Mireshghallah, Reza Shokri, and Florian Tramèr · 2022
Earlier work this paper cites.
Considerations for differentially private learning with large-scale public pretraining
Florian Tramèr, Gautam Kamath, and Nicholas Carlini · 2022
Earlier work this paper cites.
If influence functions are the answer, then what is the question?
Juhan Bae, Nathan Ng, Alston Lo, Marzyeh Ghassemi, and Roger B Grosse · 2022
Earlier work this paper cites.
Unrolling sgd: Understanding factors influencing machine unlearning
Anvith Thudi, Gabriel Deza, Varun Chandrasekaran, and Nicolas Papernot · 2022
Earlier work this paper cites.
Membership inference attacks from first principles
Nicholas Carlini, Steve Chien, Milad Nasr, Shuang Song, Andreas Terzis, and Florian Tramer · 2022
Cited alongside, same era.
Who’s harry potter? approximate unlearning in llms
Ronen Eldan and Mark Russinovich · 2023
Cited alongside, same era.
In-context unlearning: Language models as few shot unlearners
Martin Pawelczyk, Seth Neel, and Himabindu Lakkaraju · 2023
Cited alongside, same era.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al · 2023
Cited alongside, same era.
https://unlearning-challenge.github.io/assets/data/Machine_Unlearning_Metric.pdf , 2023
Evaluation for the neurips machine unlearning competition · 2023
Cited alongside, same era.
A framework for few-shot language model evaluation, 12 2023
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou · 2023
Later among the works it cites.
Detecting pretraining data from large language models
Weijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang, Daogao Liu, Terra Blevins, Danqi Chen, and Luke Zettlemoyer · 2023
Later among the works it cites.
Calibrated language models must hallucinate
Adam Tauman Kalai and Santosh S Vempala · 2023
Later among the works it cites.
Chatgpt and physicians’ malpractice risk
Michelle M Mello and Neel Guha · 2023
Later among the works it cites.
Executive order on the safe, secure, and trustworthy development and use of artificial intelligence
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Knowledge unlearning for llms: Tasks, methods, and challenges
Nianwen Si, Hao Zhang, Heyu Chang, Wenlin Zhang, Dan Qu, and Weiqiang Zhang · 2023
Cited alongside, same era.
Right to be forgotten in the era of large language models: Implications, challenges, and solutions
Dawen Zhang, Pamela Finckenberg-Broman, Thong Hoang, Shidong Pan, Zhenchang Xing, Mark Staples, and Xiwei Xu · 2023
Cited alongside, same era.
Removing rlhf protections in gpt-4 via fine-tuning
Qiusi Zhan, Richard Fang, Rohan Bindu, Akul Gupta, Tatsunori Hashimoto, and Daniel Kang · 2023
Cited alongside, same era.
Self-destructing models: Increasing the costs of harmful dual uses of foundation models
Peter Henderson, Eric Mitchell, Christopher Manning, Dan Jurafsky, and Chelsea Finn · 2023
Cited alongside, same era.
Fast yet effective machine unlearning
Ayush K Tarun, Vikram S Chundawat, Murari Mandal, and Mohan Kankanhalli · 2023
Cited alongside, same era.
Machine unlearning: A survey
Heng Xu, Tianqing Zhu, Lefeng Zhang, Wanlei Zhou, and Philip S Yu · 2023
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Joseph R Biden · 2023
Later among the works it cites.
Tofu: A task of fictitious unlearning for llms
Zhili Maini, Pratyush evand Feng, Avi Schwarzschild, Zachary C Lipton, and J Zico Kolter · 2024
Closest in time.
The wmdp benchmark: Measuring and reducing malicious use with unlearning
Nathaniel Li, Alexander Pan, Anjali Gopal, Summer Yue, Daniel Berrios, Alice Gatti, Justin D Li, Ann-Kathrin Dombrowski, Shashwat Goel, Long Phan, et al · 2024
Closest in time.
Jailbreaking is best solved by definition
Taeyoun Kim, Suhas Kotha, and Aditi Raghunathan · 2024
Closest in time.
Unlearnable algorithms for in-context learning
Andrei Muresanu, Anvith Thudi, Michael R Zhang, and Nicolas Papernot · 2024
Closest in time.
Rethinking machine unlearning for large language models
Sijia Liu, Yuanshun Yao, Jinghan Jia, Stephen Casper, Nathalie Baracaldo, Peter Hase, Xiaojun Xu, Yuguang Yao, Hang Li, Kush R Varshney, et al · 2024
Closest in time.
Eight methods to evaluate robust unlearning in llms
Aengus Lynch, Phillip Guo, Aidan Ewart, Stephen Casper, and Dylan Hadfield-Menell · 2024
Closest in time.
personal communication, March 2024
Phillip Guo · 2024
Closest in time.
Negative preference optimization: From catastrophic collapse to effective unlearning
Ruiqi Zhang, Licong Lin, Yu Bai, and Song Mei · 2024
Closest in time.
Leace: Perfect linear concept erasure in closed form
Nora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell, Edward Raff, and Stella Biderman · 2024
Closest in time.
Autodefense: Multi-agent llm defense against jailbreak attacks
Yifan Zeng, Yiran Wu, Xiao Zhang, Huazheng Wang, and Qingyun Wu · 2024
Closest in time.
Prp: Propagating universal perturbations to attack large language model guard-rails
Neal Mangaokar, Ashish Hooda, Jihye Choi, Shreyas Chandrashekaran, Kassem Fawaz, Somesh Jha, and Atul Prakash · 2024
Closest in time.
Judging llm-as-a-judge with mt-bench and chatbot arena
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al · 2024
Closest in time.
Fast machine unlearning without retraining through selective synaptic dampening
Jack Foster, Stefan Schoepf, and Alexandra Brintrup · 2024
Closest in time.