Fetching the paper…
Reading the bibliography…
Large Language Models' knowledge of how to perform cyber-security attacks, create bioweapons, and manipulate humans poses risks of misuse.
On evaluating adversarial robustness, 2019a
Nicholas Carlini, Anish Athalye, Nicolas Papernot, Wieland Brendel, Jonas Rauber, Dimitris Tsipras, Ian Goodfellow, Aleksander Madry, and Alexey Kurakin · 1902
Earlier work this paper cites.
Eternal sunshine of the spotless net: Selective forgetting in deep networks, 2020a
Aditya Golatkar, Alessandro Achille, and Stefano Soatto · 1911
Earlier work this paper cites.
Aditya Golatkar, Alessandro Achille, and Stefano Soatto · 2003
Earlier work this paper cites.
Membership inference attacks against machine learning models (s&p’17)
Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov · 2016
Earlier work this paper cites.
Knowledge unlearning for mitigating privacy risks in language models, 2022
Joel Jang, Dongkeun Yoon, Sohee Yang, Sungmin Cha, Moontae Lee, Lajanugen Logeswaran, and Minjoon Seo · 2022
Earlier work this paper cites.
Continual learning and private unlearning, 2022
Bo Liu, Qiang Liu, and Peter Stone · 2022
Earlier work this paper cites.
A survey of machine unlearning
Thanh Tam Nguyen, Thanh Trung Huynh, Phi Le Nguyen, Alan Wee-Chung Liew, Hongzhi Yin, and Quoc Viet Hung Nguyen · 2022
Earlier work this paper cites.
Emergent abilities of large language models, 2022
Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus · 2022
Earlier work this paper cites.
Symbolic discovery of optimization algorithms, 2023
Xiangning Chen, Chen Liang, Da Huang, Esteban Real, Kaiyuan Wang, Yao Liu, Hieu Pham, Xuanyi Dong, Thang Luong, Cho-Jui Hsieh, Yifeng Lu, and Quoc V. Le · 2023
Earlier work this paper cites.
Deep reinforcement learning from human preferences, 2023
Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei · 2023
Earlier work this paper cites.
Who’s harry potter? approximate unlearning in llms, 2023
Ronen Eldan and Mark Russinovich · 2023
Earlier work this paper cites.
Self-destructing models: Increasing the costs of harmful dual uses of foundation models, 2023
Peter Henderson, Eric Mitchell, Christopher D. Manning, Dan Jurafsky, and Chelsea Finn · 2023
Cited alongside, same era.
Fine-tuning aligned language models compromises safety, even when users do not intend to!, 2023
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson · 2023
Cited alongside, same era.
A conversation with bing’s chatbot left me deeply unsettled
Kevin Roose · 2023
Cited alongside, same era.
Knowledge unlearning for llms: Tasks, methods, and challenges, 2023
Nianwen Si, Hao Zhang, Heyu Chang, Wenlin Zhang, Dan Qu, and Weiqiang Zhang · 2023
Cited alongside, same era.
Fast yet effective machine unlearning
Ayush K Tarun, Vikram S Chundawat, Murari Mandal, and Mohan Kankanhalli · 2023
Intrinsic evaluation of unlearning using parametric knowledge traces, 2024
Yihuai Hong, Lei Yu, Haiqin Yang, Shauli Ravfogel, and Mor Geva · 2024
Closest in time.
Jogging the memory of unlearned llms through targeted relearning attacks, 2024
Shengyuan Hu, Yiwei Fu, Zhiwei Steven Wu, and Virginia Smith · 2024
Closest in time.
Mechanistically analyzing the effects of fine-tuning on procedurally defined tasks, 2024
Samyak Jain, Robert Kirk, Ekdeep Singh Lubana, Robert P. Dick, Hidenori Tanaka, Edward Grefenstette, Tim Rocktäschel, and David Scott Krueger · 2024
Closest in time.
The llama 3 herd of models
AI@Meta Llama Team · 2024
Closest in time.
Tofu: A task of fictitious unlearning for llms, 2024
Pratyush Maini, Zhili Feng, Avi Schwarzschild, Zachary C. Lipton, and J. Zico Kolter · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Zephyr: Direct distillation of lm alignment
Lewis Tunstall, Edward Beeching, Nathan Lambert, Nazneen Rajani, Kashif Rasul, Younes Belkada, Shengyi Huang, Leandro von Werra, Clémentine Fourrier, Nathan Habib, et al · 2023
Cited alongside, same era.
Jailbroken: How does llm safety training fail?, 2023
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt · 2023
Cited alongside, same era.
Executive order on the safe, secure, and trustworthy development and use of artificial intelligence, 2023
White House · 2023
Cited alongside, same era.
Universal and transferable adversarial attacks on aligned language models, 2023
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J. Zico Kolter, and Matt Fredrikson · 2023
Cited alongside, same era.
Many-shot jailbreaking
Cem Anil, Esin Durmus, Mrinank Sharma, Joe Benton, Sandipan Kundu, Joshua Batson, Nina Rimsky, Meg Tong, Jesse Mu, Daniel Ford, et al · 2024
Cited alongside, same era.
Llm agents can autonomously hack websites, 2024
Richard Fang, Rohan Bindu, Akul Gupta, Qiusi Zhan, and Daniel Kang · 2024
Cited alongside, same era.
The secret sharer: Evaluating and testing unintended memorization in neural networks, 2019b
Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, and Dawn Song
Cited in the paper.
Closest in time.
Hello gpt-4o
OpenAI · 2024
Closest in time.
The fineweb datasets: Decanting the web for the finest text data at scale, 2024
Guilherme Penedo, Hynek Kydlíček, Loubna Ben allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, and Thomas Wolf · 2024
Closest in time.
Representation noising effectively prevents harmful fine-tuning on llms, 2024
Domenic Rosati, Jan Wehner, Kai Williams, Łukasz Bartoszcze, David Atanasov, Robie Gonzales, Subhabrata Majumdar, Carsten Maple, Hassan Sajjad, and Frank Rudzicz · 2024
Closest in time.
Toward robust unlearning for LLMs
Rishub Tamirisa, Bhrugu Bharathi, Andy Zhou, Bo Li, and Mantas Mazeika · 2024
Closest in time.
An adversarial perspective on machine unlearning for ai safety, 2024
Jakub Łucki, Boyi Wei, Yangsibo Huang, Peter Henderson, Florian Tramèr, and Javier Rando · 2024
Closest in time.
Open problems in machine unlearning for ai safety, 2025
Fazl Barez, Tingchen Fu, Ameya Prabhu, Stephen Casper, Amartya Sanyal, Adel Bibi, Aidan O’Gara, Robert Kirk, Ben Bucknall, Tim Fist, Luke Ong, Philip Torr, Kwok-Yan Lam, Robert Trager, David Krueger, Sören Mindermann, José Hernandez-Orallo, Mor Geva, and Yarin Gal · 2025
Closest in time.