Fetching the paper…
Reading the bibliography…
Machine unlearning has been used to remove unwanted knowledge acquired by large language models (LLMs).
Maximum likelihood from incomplete data via the EM algorithm
A. P. Dempster, N. M. Laird, and D. B. Rubin · 1977
Earlier work this paper cites.
The concave-convex procedure (cccp)
Alan L Yuille and Anand Rangarajan · 2001
Earlier work this paper cites.
Convex optimization
Stephen Boyd and Lieven Vandenberghe · 2004
Earlier work this paper cites.
Learning multiple layers of features from tiny images
Alex Krizhevsky, Geoffrey Hinton, et al · 2009
Earlier work this paper cites.
Multiple-gradient descent algorithm (mgda) for multiobjective optimization
Jean-Antoine Désidéri · 2012
Earlier work this paper cites.
Unified expectation maximization
Rajhans Samdani, Ming-Wei Chang, and Dan Roth · 2012
Earlier work this paper cites.
Deep residual learning for image recognition, 2015
Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun · 2015
Earlier work this paper cites.
Gradnorm: Gradient normalization for adaptive loss balancing in deep multitask networks
Zhao Chen, Vijay Badrinarayanan, Chen-Yu Lee, and Andrew Rabinovich · 2018
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al · 2019
Earlier work this paper cites.
Just pick a sign: Optimizing deep multitask models with gradient sign dropout, 2020
Zhao Chen, Jiquan Ngiam, Yanping Huang, Thang Luong, Henrik Kretzschmar, Yuning Chai, and Dragomir Anguelov · 2020
Earlier work this paper cites.
Gradient vaccine: Investigating and improving multi-task optimization in massively multilingual models, 2020
Zirui Wang, Yulia Tsvetkov, Orhan Firat, and Yuan Cao · 2020
Earlier work this paper cites.
Gradient surgery for multi-task learning
Tianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine, Karol Hausman, and Chelsea Finn · 2020
Earlier work this paper cites.
Machine unlearning
Lucas Bourtoule, Varun Chandrasekaran, Christopher A Choquette-Choo, Hengrui Jia, Adelin Travers, Baiwu Zhang, David Lie, and Nicolas Papernot · 2021
Earlier work this paper cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al · 2021
Earlier work this paper cites.
Multi-task learning in natural language processing: An overview
Shijie Chen, Yu Zhang, and Qiang Yang · 2021
Earlier work this paper cites.
Editing factual knowledge in language models
Nicola De Cao, Wilker Aziz, and Ivan Titov · 2021
Earlier work this paper cites.
Adaptive machine unlearning
Varun Gupta, Christopher Jung, Seth Neel, Aaron Roth, Saeed Sharifi-Malvajerdi, and Chris Waites · 2021
Earlier work this paper cites.
Lora: Low-rank adaptation of large language models
Edward J Hu, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen, et al · 2021
Earlier work this paper cites.
Reasonable effectiveness of random weighting: A litmus test for multi-task learning
Baijiong Lin, Feiyang Ye, Yu Zhang, and Ivor Tsang · 2021
Earlier work this paper cites.
Towards impartial multi-task learning
Liyang Liu, Yi Li, Zhanghui Kuang, J Xue, Yimin Chen, Wenming Yang, Qingmin Liao, and Wayne Zhang · 2021
Cited alongside, same era.
Machine unlearning via algorithmic stability
Enayat Ullah, Tung Mai, Anup Rao, Ryan A Rossi, and Raman Arora · 2021
Cited alongside, same era.
Concealed data poisoning attacks on NLP models
Eric Wallace, Tony Zhao, Shi Feng, and Sameer Singh · 2021
Cited alongside, same era.
Rotograd: Gradient homogenization in multitask learning, 2022
Adrián Javaloy and Isabel Valera · 2022
Cited alongside, same era.
Continual learning and private unlearning
Bo Liu, Qiang Liu, and Peter Stone · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback, 2022
Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe · 2022
Knowledge unlearning for llms: Tasks, methods, and challenges, 2023
Nianwen Si, Hao Zhang, Heyu Chang, Wenlin Zhang, Dan Qu, and Weiqiang Zhang · 2023
Later among the works it cites.
Exhibit j, 2023
The New York Times · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al · 2023
Later among the works it cites.
A survey of multi-task learning in natural language processing: Regarding task relatedness and training methods
Zhihan Zhang, Wenhao Yu, Mengxia Yu, Zhichun Guo, and Meng Jiang · 2023
Later among the works it cites.
Gradient descent with generalized newton’s method
Zhiqi Bu and Shiyun Xu · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Do current multi-task optimization methods in deep learning even help?
Derrick Xin, Behrooz Ghorbani, Justin Gilmer, Ankush Garg, and Orhan Firat · 2022
Cited alongside, same era.
Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al · 2023
Cited alongside, same era.
Unlearn what you want to forget: Efficient unlearning for LLMs
Jiaao Chen and Diyi Yang · 2023
Cited alongside, same era.
On the pareto front of multilingual neural machine translation
Liang Chen, Shuming Ma, Dongdong Zhang, Furu Wei, and Baobao Chang · 2023
Cited alongside, same era.
Learning-rate-free learning by d-adaptation, 2023
Aaron Defazio and Konstantin Mishchenko · 2023
Cited alongside, same era.
Who’s harry potter? approximate unlearning in llms
Ronen Eldan and Mark Russinovich · 2023
Cited alongside, same era.
Kongyang Chen, Zixin Wang, Bing Mi, Waixi Liu, Shaowei Wang, Xiaojun Ren, and Jiaxing Shen · 2024
Closest in time.
Dowg unleashed: An efficient universal parameter-free gradient descent method, 2024
Ahmed Khaled, Konstantin Mishchenko, and Chi Jin · 2024
Closest in time.
Backdoorllm: A comprehensive benchmark for backdoor attacks on large language models, 2024
Yige Li, Hanxun Huang, Yunhan Zhao, Xingjun Ma, and Jun Sun · 2024
Closest in time.
Conflict-averse gradient descent for multi-task learning, 2024
Bo Liu, Xingchao Liu, Xiaojie Jin, Peter Stone, and Qiang Liu · 2024
Closest in time.
Rethinking machine unlearning for large language models
Sijia Liu, Yuanshun Yao, Jinghan Jia, Stephen Casper, Nathalie Baracaldo, Peter Hase, Xiaojun Xu, Yuguang Yao, Hang Li, Kush R Varshney, et al · 2024
Closest in time.
Rethinking machine unlearning for large language models, 2024
Sijia Liu, Yuanshun Yao, Jinghan Jia, Stephen Casper, Nathalie Baracaldo, Peter Hase, Yuguang Yao, Chris Yuhao Liu, Xiaojun Xu, Hang Li, Kush R. Varshney, Mohit Bansal, Sanmi Koyejo, and Yang Liu · 2024
Closest in time.
Tofu: A task of fictitious unlearning for llms
Pratyush Maini, Zhili Feng, Avi Schwarzschild, Zachary C Lipton, and J Zico Kolter · 2024
Closest in time.
Prodigy: An expeditiously adaptive parameter-free learner, 2024
Konstantin Mishchenko and Aaron Defazio · 2024
Closest in time.
In-context unlearning: Language models as few shot unlearners
Martin Pawelczyk, Seth Neel, and Himabindu Lakkaraju · 2024
Closest in time.
Muse: Machine unlearning six-way evaluation for language models
Weijia Shi, Jaechan Lee, Yangsibo Huang, Sadhika Malladi, Jieyu Zhao, Ari Holtzman, Daogao Liu, Luke Zettlemoyer, Noah A Smith, and Chiyuan Zhang · 2024
Closest in time.
Akew: Assessing knowledge editing in the wild, 2024
Xiaobao Wu, Liangming Pan, William Yang Wang, and Anh Tuan Luu · 2024
Closest in time.
Right to be forgotten in the era of large language models: Implications, challenges, and solutions, 2024
Dawen Zhang, Pamela Finckenberg-Broman, Thong Hoang, Shidong Pan, Zhenchang Xing, Mark Staples, and Xiwei Xu · 2024
Closest in time.
Negative preference optimization: From catastrophic collapse to effective unlearning
Ruiqi Zhang, Licong Lin, Yu Bai, and Song Mei · 2024
Closest in time.