Fetching the paper…
Reading the bibliography…
Information removal or suppression in large language models (LLMs) is a desired functionality, useful in AI regulation, legal compliance, safety, and privacy.
Catastrophic interference in connectionist networks: The sequential learning problem
Michael McCloskey and Neal J. Cohen · 1989
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Language models are few-shot learners, 2020
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, et al · 2005
Earlier work this paper cites.
Descent-to-delete: Gradient-based methods for machine unlearning, 2020
Seth Neel, Aaron Roth, and Saeed Sharifi-Malvajerdi · 2007
Earlier work this paper cites.
Better fine-tuning by reducing representational collapse, 2020
Armen Aghajanyan, Akshat Shrivastava, Anchit Gupta, Naman Goyal, Luke Zettlemoyer, and Sonal Gupta · 2008
Earlier work this paper cites.
Measuring massive multitask language understanding, 2021
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2009
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay · 2011
Earlier work this paper cites.
General data protection regulation (gdpr), 2016
European Union · 2016
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models, 2021
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2021
Earlier work this paper cites.
Ccpa regulations: Final regulation text, 2021
California Department of Justice OAG · 2021
Earlier work this paper cites.
Knowledge unlearning for mitigating privacy risks in language models
Joel Jang, Dongkeun Yoon, Sohee Yang, Sungmin Cha, Moontae Lee, Lajanugen Logeswaran, and Minjoon Seo · 2022
Earlier work this paper cites.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Earlier work this paper cites.
The falcon series of open language models, 2023
Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, et al · 2023
Earlier work this paper cites.
Jinze Bai, Shuai Bai, Yunfei Chu, Zeyu Cui, Kai Dang, Xiaodong Deng, Yang Fan, Wenbin Ge, Yu Han, Fei Huang, Binyuan Hui, Luo Ji, Mei Li, Junyang Lin, Runji Lin, Dayiheng Liu, Gao Liu, Chengqiang Lu, Keming Lu, Jianxin Ma, Rui Men, Xingzhang Ren, Xuancheng Ren, Chuanqi Tan, Sinan Tan, Jianhong Tu, Peng Wang, Shijie Wang, Wei Wang, Shengguang Wu, Benfeng Xu, Jin Xu, An Yang, Hao Yang, Jian Yang, Shusheng Yang, Yang Yao, Bowen Yu, Hongyi Yuan, Zheng Yuan, Jianwei Zhang, Xingxuan Zhang, Yichang Zhang, Zhenru Zhang, Chang Zhou, Jingren Zhou, Xiaohuan Zhou, and Tianhang Zhu · 2023
Earlier work this paper cites.
Chateval: Towards better llm-based evaluators through multi-agent debate, 2023
Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, Jie Fu, and Zhiyuan Liu · 2023
Earlier work this paper cites.
Zero-shot machine unlearning
Vikram S. Chundawat, Ayush K. Tarun, Murari Mandal, and Mohan Kankanhalli · 2023
Earlier work this paper cites.
Who’s harry potter? approximate unlearning in llms, 2023
Ronen Eldan and Mark Russinovich · 2023
Earlier work this paper cites.
Fast machine unlearning without retraining through selective synaptic dampening, 2023
Jack Foster, Stefan Schoepf, and Alexandra Brintrup · 2023
Earlier work this paper cites.
Foundation models and fair use, 2023
Peter Henderson, Xuechen Li, Dan Jurafsky, Tatsunori Hashimoto, Mark A. Lemley, and Percy Liang · 2023
Earlier work this paper cites.
Preventing verbatim memorization in language models gives a false sense of privacy, 2023
Daphne Ippolito, Florian Tramèr, Milad Nasr, Chiyuan Zhang, Matthew Jagielski, Katherine Lee, Christopher A. Choquette-Choo, and Nicholas Carlini · 2023
Earlier work this paper cites.
Copyright violations and large language models, 2023
Antonia Karamolegkou, Jiaang Li, Li Zhou, and Anders Søgaard · 2023
Earlier work this paper cites.
Digital personal data protection act, 2023, 2023
Indian Legislative · 2023
Cited alongside, same era.
Let’s verify step by step, 2023
Hunter Lightman, Vineet Kosaraju, Yura Burda, Harri Edwards, Bowen Baker, Teddy Lee, Jan Leike, John Schulman, Ilya Sutskever, and Karl Cobbe · 2023
Cited alongside, same era.
In-context unlearning: Language models as few shot unlearners
Martin Pawelczyk, Seth Neel, and Himabindu Lakkaraju · 2023
Cited alongside, same era.
Snap: Self-supervised neural maps for visual positioning and semantic understanding, 2023
Paul-Edouard Sarlin, Eduard Trulls, Marc Pollefeys, Jan Hosang, and Simon Lynen · 2023
Cited alongside, same era.
Scalable and transferable black-box jailbreaks for language models via persona modulation, 2023
Eight methods to evaluate robust unlearning in llms
Aengus Lynch, Phillip Guo, Aidan Ewart, Stephen Casper, and Dylan Hadfield-Menell · 2024
Later among the works it cites.
Tofu: A task of fictitious unlearning for llms
Pratyush Maini, Zhili Feng, Avi Schwarzschild, Zachary C Lipton, and J Zico Kolter · 2024
Later among the works it cites.
Prp: Propagating universal perturbations to attack large language model guard-rails, 2024
Neal Mangaokar, Ashish Hooda, Jihye Choi, Shreyas Chandrashekaran, Kassem Fawaz, Somesh Jha, and Atul Prakash · 2024
Later among the works it cites.
Unlearnable algorithms for in-context learning, 2024
Andrei Muresanu, Anvith Thudi, Michael R. Zhang, and Nicolas Papernot · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Rusheb Shah, Quentin Feuillade-Montixi, Soroush Pour, Arush Tagade, Stephen Casper, and Javier Rando · 2023
Cited alongside, same era.
Fast yet effective machine unlearning
Ayush K. Tarun, Vikram S. Chundawat, Murari Mandal, and Mohan Kankanhalli · 2023
Cited alongside, same era.
Llama 2: Open foundation and fine-tuned chat models, 2023
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Cited alongside, same era.
Chain-of-thought prompting elicits reasoning in large language models, 2023
Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed Chi, Quoc Le, and Denny Zhou · 2023
Cited alongside, same era.
Fedrecovery: Differentially private machine unlearning for federated learning frameworks
Lefeng Zhang, Tianqing Zhu, Haibin Zhang, Ping Xiong, and Wanlei Zhou · 2023
Cited alongside, same era.
Many-shot jailbreaking
Cem Anil, Esin Durmus, Mrinank Sharma, Joe Benton, Sandipan Kundu, Joshua Batson, Nina Rimsky, Meg Tong, Jesse Mu, Daniel Ford, et al · 2024
Cited alongside, same era.
Internlm2 technical report, 2024
Zheng Cai, Maosong Cao, Haojiong Chen, Kai Chen, et al · 2024
Cited alongside, same era.
Opt-out: Investigating entity-level unlearning for large language models via optimal transport, 2024
Minseok Choi, Daniel Rim, Dohyun Lee, and Jaegul Choo · 2024
Cited alongside, same era.
OpenAI · 2024
Later among the works it cites.
Jiahao Qiu, Yifu Lu, Yifan Zeng, Jiacheng Guo, Jiayi Geng, Huazheng Wang, Kaixuan Huang, Yue Wu, and Mengdi Wang · 2024
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model, 2024
Rafael Rafailov, Archit Sharma, Eric Mitchell, Stefano Ermon, Christopher D. Manning, and Chelsea Finn · 2024
Later among the works it cites.
Tricking llms into disobedience: Formalizing, analyzing, and detecting jailbreaks, 2024
Abhinav Rao, Sachin Vashistha, Atharva Naik, Somak Aditya, and Monojit Choudhury · 2024
Later among the works it cites.
Leo Schwinn, David Dobre, Sophie Xhonneux, Gauthier Gidel, and Stephan Gunnemann · 2024
Later among the works it cites.
Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang · 2024
Later among the works it cites.
Unstar: Unlearning with self-taught anti-sample reasoning for llms, 2024
Yash Sinha, Murari Mandal, and Mohan Kankanhalli · 2024
Later among the works it cites.
Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024
Charlie Snell, Jaehoon Lee, Kelvin Xu, and Aviral Kumar · 2024
Later among the works it cites.
Beyond memorization: Violating privacy via inference with large language models, 2024
Robin Staab, Mark Vero, Mislav Balunović, and Martin Vechev · 2024
Later among the works it cites.
Qwen2.5: A party of foundation models, September 2024
Qwen Team · 2024
Later among the works it cites.
Guardrail baselines for unlearning in llms
Pratiksha Thaker, Yash Maurya, Shengyuan Hu, Zhiwei Steven Wu, and Virginia Smith · 2024
Later among the works it cites.
Evaluating deep unlearning in large language models
Ruihan Wu, Chhavi Yadav, Russ Salakhutdinov, and Kamalika Chaudhuri · 2024
Later among the works it cites.
Machine unlearning: Solutions and challenges
Jie Xu, Zihan Wu, Cong Wang, and Xiaohua Jia · 2024
Later among the works it cites.
Large language model unlearning, 2024
Yuanshun Yao, Xiaojun Xu, and Yang Liu · 2024
Later among the works it cites.
Negative preference optimization: From catastrophic collapse to effective unlearning, 2024
Ruiqi Zhang, Licong Lin, Yu Bai, and Song Mei · 2024
Later among the works it cites.
Frequently occurring surnames from the 2010 census
U.S. Census Bureau · 2025
Closest in time.