Fetching the paper…
Reading the bibliography…
This study investigates the machine unlearning techniques within the context of large language models (LLMs), referred to as \textit{LLM unlearning}.
General data protection regulation (GDPR)
Council of the European Union. 2016 · 2016
Earlier work this paper cites.
Proximal policy optimization algorithms
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. 2017 · 2017
Earlier work this paper cites.
Continual learning and private unlearning
Bo Liu, Qiang Liu, and Peter Stone. 2022 · 2022
Earlier work this paper cites.
Quark: Controllable text generation with reinforced unlearning
Ximing Lu, Sean Welleck, Jack Hessel, and Yejin Choi. 2022 · 2022
Earlier work this paper cites.
Editing models with task arithmetic
Gabriel Ilharco, Marco Tulio Ribeiro, Mitchell Wortsman, Ludwig Schmidt, Hannaneh Hajishirzi, and Ali Farhadi. 2023 · 2023
Earlier work this paper cites.
Knowledge unlearning for mitigating privacy risks in language models
Joel Jang, Dongkeun Yoon, Sohee Yang, Sungmin Cha, Moontae Lee, Lajanugen Logeswaran, and Minjoon Seo. 2023 · 2023
Earlier work this paper cites.
Preserving privacy through dememorization: An unlearning technique for mitigating memorization risks in language models
Aly Kassem, Omar Mahmoud, and Sherif Saad. 2023 · 2023
Earlier work this paper cites.
Direct preference optimization: Your language model is secretly a reward model
Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn. 2023 · 2023
Earlier work this paper cites.
Knowledge unlearning for llms: Tasks, methods, and challenges
Nianwen Si, Hao Zhang, Heyu Chang, Wenlin Zhang, Dan Qu, and Weiqiang Zhang. 2023 · 2023
Earlier work this paper cites.
DEPN: Detecting and editing privacy neurons in pretrained language models
Xinwei Wu, Junzhuo Li, Minghui Xu, Weilong Dong, Shuangzhi Wu, Chao Bian, and Deyi Xiong. 2023 · 2023
Earlier work this paper cites.
Refusal in language models is mediated by a single direction
Andy Arditi, Oscar Balcells Obeso, Aaquib Syed, Daniel Paleka, Nina Rimsky, Wes Gurnee, and Neel Nanda. 2024 · 2024
Earlier work this paper cites.
Opt-out: Investigating entity-level unlearning for large language models via optimal transport
Minseok Choi, Daniel Rim, Dohyun Lee, and Jaegul Choo. 2024 · 2024
Earlier work this paper cites.
Laying down harmonised rules on artificial intelligence (artificial intelligence act) and amending certain union legislative acts
Council of the European Union. 2024 · 2024
Earlier work this paper cites.
Unmemorization in large language models via self-distillation and deliberate imagination
Yijiang River Dong, Hongzhou Lin, Mikhail Belkin, Ramon Huerta, and Ivan Vulić. 2024 · 2024
Earlier work this paper cites.
Clear: Character unlearning in textual and visual modalities
Alexey Dontsov, Dmitrii Korzh, Alexey Zhavoronkin, Boris Mikheev, Denis Bobkov, Aibek Alanov, Oleg Y Rogov, Ivan Oseledets, and Elena Tutubalina. 2024 · 2024
Cited alongside, same era.
Who’s harry potter? approximate unlearning for LLMs
Ronen Eldan, Mark Russinovich, and Mark Russinovich. 2024 · 2024
Cited alongside, same era.
Mechanistic unlearning: Robust knowledge unlearning and editing via mechanistic localization
Phillip Guo, Aaquib Syed, Abhay Sheshadri, Aidan Ewart, and Gintare Karolina Dziugaite. 2024 · 2024
Cited alongside, same era.
Intrinsic evaluation of unlearning using parametric knowledge traces
Yihuai Hong, Lei Yu, Haiqin Yang, Shauli Ravfogel, and Mor Geva. 2024 · 2024
Cited alongside, same era.
Jogging the memory of unlearned llms through targeted relearning attacks
Benchmarking vision language model unlearning via fictitious facial identity dataset
Yingzi Ma, Jiongxiao Wang, Fei Wang, Siyuan Ma, Jiazhao Li, Xiujun Li, Furong Huang, Lichao Sun, Bo Li, Yejin Choi, and 1 others. 2024 · 2024
Later among the works it cites.
TOFU: A task of fictitious unlearning for LLMs
Pratyush Maini, Zhili Feng, Avi Schwarzschild, Zachary Chase Lipton, and J Zico Kolter. 2024 · 2024
Later among the works it cites.
Extracting unlearned information from llms with activation steering
Atakan Seyitoğlu, Aleksei Kuvshinov, Leo Schwinn, and Stephan Günnemann. 2024 · 2024
Later among the works it cites.
Guardrail baselines for unlearning in llms
Pratiksha Thaker, Yash Maurya, Shengyuan Hu, Zhiwei Steven Wu, and Virginia Smith. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Shengyuan Hu, Yiwei Fu, Steven Wu, and Virginia Smith. 2024 · 2024
Cited alongside, same era.
An information theoretic metric for evaluating unlearning models
Dongjae Jeon, Wonje Jeung, Taeheon Kim, Albert No, and Jonghyun Choi. 2024 · 2024
Cited alongside, same era.
Reversing the forget-retain objectives: An efficient llm unlearning framework from logit difference
Jiabao Ji, Yujian Liu, Yang Zhang, Gaowen Liu, Ramana Rao Kompella, Sijia Liu, and Shiyu Chang. 2024 · 2024
Cited alongside, same era.
WAGLE: Strategic weight attribution for effective and modular unlearning in large language models
Jinghan Jia, Jiancheng Liu, Yihua Zhang, Parikshit Ram, Nathalie Baracaldo, and Sijia Liu. 2024 · 2024
Cited alongside, same era.
Rwku: Benchmarking real-world knowledge unlearning for large language models
Zhuoran Jin, Pengfei Cao, Chenhao Wang, Zhitao He, Hongbang Yuan, Jiachun Li, Yubo Chen, Kang Liu, and Jun Zhao. 2024 · 2024
Cited alongside, same era.
Single image unlearning: Efficient machine unlearning in multimodal large language models
Jiaqi Li, Qianshan Wei, Chuanyi Zhang, Guilin Qi, Miaozeng Du, Yongrui Chen, Sheng Bi, and Fan Liu. 2024a · 2024
Cited alongside, same era.
Towards safer large language models through machine unlearning
Zheyuan Liu, Guangyao Dou, Zhaoxuan Tan, Yijun Tian, and Meng Jiang. 2024d · 2024
Cited alongside, same era.
Towards transfer unlearning: Empirical evidence of cross-domain bias mitigation
Huimin Lu, Masaru Isonuma, Junichiro Mori, and Ichiro Sakata. 2024 · 2024
Cited alongside, same era.
Bichen Wang, Yuzhe Zi, Yixin Sun, Yanyan Zhao, and Bing Qin. 2024 · 2024
Later among the works it cites.
Machine unlearning for traditional models and large language models: A short survey
Yi Xu. 2024 · 2024
Later among the works it cites.
Machine unlearning of pre-trained large language models
Jin Yao, Eli Chien, Minxin Du, Xinyao Niu, Tianhao Wang, Zezhou Cheng, and Xiang Yue. 2024 · 2024
Later among the works it cites.
A closer look at machine unlearning for large language models
Xiaojian Yuan, Tianyu Pang, Chao Du, Kejiang Chen, Weiming Zhang, and Min Lin. 2024 · 2024
Later among the works it cites.
On effects of steering latent representation for large language model unlearning
Huu-Tien Dang, Tin Pham, Hoang Thanh-Tung, and Naoya Inoue. 2025 · 2025
Closest in time.
Rethinking machine unlearning for large language models
Sijia Liu, Yuanshun Yao, Jinghan Jia, Stephen Casper, Nathalie Baracaldo, Peter Hase, Yuguang Yao, Chris Yuhao Liu, Xiaojun Xu, Hang Li, and 1 others. 2025 · 2025
Closest in time.
In-context unlearning: language models as few-shot unlearners
Martin Pawelczyk, Seth Neel, and Himabindu Lakkaraju. 2025 · 2025
Closest in time.
Catastrophic failure of LLM unlearning via quantization
Zhiwei Zhang, Fali Wang, Xiaomin Li, Zongyu Wu, Xianfeng Tang, Hui Liu, Qi He, Wenpeng Yin, and Suhang Wang. 2025 · 2025
Closest in time.
Ethos: Rectifying language models in orthogonal parameter space
Lei Gao, Yue Niu, Tingting Tang, Salman Avestimehr, and Murali Annavaram. 2024 · 2068
Closest in time.