Fetching the paper…
Reading the bibliography…
Representation Misdirection for Unlearning (RMU), which steers model representation in the intermediate layer to a target random representation, is an effective method for large language model (LLM) unlearning.
Towards Making Systems Forget with Machine Unlearning
Cao, Y.; and Yang, J. 2015 · 2015
Earlier work this paper cites.
Approximate data deletion from machine learning models
Izzo, Z.; Smart, M. A.; Chaudhuri, K.; and Zou, J. 2021 · 2016
Earlier work this paper cites.
A Baseline for Detecting Misclassified and Out-of-Distribution Examples in Neural Networks
Hendrycks, D.; and Gimpel, K. 2017 · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Koh, P. W.; and Liang, P. 2017 · 2017
Earlier work this paper cites.
Making ai forget you: Data deletion in machine learning
Ginart, A.; Guan, M.; Valiant, G.; and Zou, J. Y. 2019 · 2019
Earlier work this paper cites.
Decoupled Weight Decay Regularization
Loshchilov, I.; and Hutter, F. 2019 · 2019
Earlier work this paper cites.
Eternal sunshine of the spotless net: Selective forgetting in deep networks
Golatkar, A.; Achille, A.; and Soatto, S. 2020 · 2020
Earlier work this paper cites.
Energy-based out-of-distribution detection
Liu, W.; Wang, X.; Owens, J.; and Li, Y. 2020 · 2020
Earlier work this paper cites.
Machine unlearning
Bourtoule, L.; Chandrasekaran, V.; Choquette-Choo, C. A.; Jia, H.; Travers, A.; Zhang, B.; Lie, D.; and Papernot, N. 2021 · 2021
Earlier work this paper cites.
Measuring Massive Multitask Language Understanding
Hendrycks, D.; Burns, C.; Basart, S.; Zou, A.; Mazeika, M.; Song, D.; and Steinhardt, J. 2021 · 2021
Earlier work this paper cites.
Confident learning: Estimating uncertainty in dataset labels
Northcutt, C.; Jiang, L.; and Chuang, I. 2021 · 2021
Earlier work this paper cites.
Remember what you want to forget: Algorithms for machine unlearning
Sekhari, A.; Acharya, J.; Kamath, G.; and Suresh, A. T. 2021 · 2021
Earlier work this paper cites.
Machine unlearning of features and labels
Warnecke, A.; Pirch, L.; Wressnegger, C.; and Rieck, K. 2021 · 2021
Earlier work this paper cites.
Graph unlearning
Chen, M.; Zhang, Z.; Wang, T.; Backes, M.; Humbert, M.; and Zhang, Y. 2022 · 2022
Earlier work this paper cites.
Federated Unlearning: How to Efficiently Erase a Client in FL?
Halimi, A.; Kadhe, S. R.; Rawat, A.; and Angel, N. B. 2022 · 2022
Earlier work this paper cites.
Scaling Out-of-Distribution Detection for Real-World Settings
Hendrycks, D.; Basart, S.; Mazeika, M.; Zou, A.; Kwon, J.; Mostajabi, M.; Steinhardt, J.; and Song, D. 2022 · 2022
Earlier work this paper cites.
Learn to forget: Machine unlearning via neuron masking
Ma, Z.; Liu, Y.; Liu, X.; Liu, J.; Ma, J.; and Ren, K. 2022 · 2022
Earlier work this paper cites.
Pointer Sentinel Mixture Models
Merity, S.; Xiong, C.; Bradbury, J.; and Socher, R. 2022 · 2022
Earlier work this paper cites.
A survey of machine unlearning
Nguyen, T. T.; Huynh, T. T.; Nguyen, P. L.; Liew, A. W.-C.; Yin, H.; and Nguyen, Q. V. H. 2022 · 2022
Earlier work this paper cites.
Out-of-distribution detection with deep nearest neighbors
Sun, Y.; Ming, Y.; Zhu, X.; and Li, Y. 2022 · 2022
Earlier work this paper cites.
Unrolling sgd: Understanding factors influencing machine unlearning
Thudi, A.; Deza, G.; Chandrasekaran, V.; and Papernot, N. 2022 · 2022
Earlier work this paper cites.
Federated Unlearning via Class-Discriminative Pruning
Wang, J.; Guo, S.; Xie, X.; and Qi, H. 2022 · 2022
Earlier work this paper cites.
Mitigating neural network overconfidence with logit normalization
Wei, H.; Xie, R.; Cheng, H.; Feng, L.; An, B.; and Li, Y. 2022 · 2022
Earlier work this paper cites.
LEACE: Perfect linear concept erasure in closed form
Belrose, N.; Schneider-Joseph, D.; Ravfogel, S.; Cotterell, R.; Raff, E.; and Biderman, S. 2023 · 2023
Earlier work this paper cites.
Fast federated machine unlearning with nonlinear functional theory
Che, T.; Zhou, Y.; Zhang, Z.; Lyu, L.; Liu, J.; Yan, D.; Dou, D.; and Huan, J. 2023 · 2023
Cited alongside, same era.
Cheng, J.; and Amiri, H. 2023 · 2023
Cited alongside, same era.
GNNDelete: A General Strategy for Unlearning in Graph Neural Networks
Cheng, J.; Dasoulas, G.; He, H.; Agarwal, C.; and Zitnik, M. 2023 · 2023
Cited alongside, same era.
Efficient Model Updates for Approximate Unlearning of Graph-Structured Data
Chien, E.; Pan, C.; and Milenkovic, O. 2023 · 2023
Cited alongside, same era.
Choi, D.; and Na, D. 2023 · 2023
Cited alongside, same era.
Erasing Undesirable Concepts in Diffusion Models with Adversarial Preservation
Bui, T.-A.; Long, V.; Doan, K.; Le, T.; Montague, P.; Abraham, T.; and Phung, D. 2024 · 2024
Closest in time.
Learning to unlearn: Instance-wise unlearning for pre-trained classifiers
Cha, S.; Cho, S.; Hwang, D.; Lee, H.; Moon, T.; and Lee, M. 2024 · 2024
Closest in time.
Post-Training Attribute Unlearning in Recommender Systems
Chen, C.; Zhang, Y.; Li, Y.; Wang, J.; Qi, L.; Xu, X.; Zheng, X.; and Yin, J. 2024 · 2024
Closest in time.
Cooper, A. F.; Choquette-Choo, C. A.; Bogen, M.; Jagielski, M.; Filippova, K.; Liu, K. Z.; Chouldechova, A.; Hayes, J.; Huang, Y.; Mireshghallah, N.; et al. 2024 · 2024
Closest in time.
SalUn: Empowering Machine Unlearning via Gradient-based Weight Saliency in Both Image Classification and Generation
Fan, C.; Liu, J.; Zhang, Y.; Wong, E.; Wei, D.; and Liu, S. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Safe: Machine unlearning with shard graphs
Dukler, Y.; Bowman, B.; Achille, A.; Golatkar, A.; Swaminathan, A.; and Soatto, S. 2023 · 2023
Cited alongside, same era.
Who’s Harry Potter? Approximate Unlearning in LLMs
Eldan, R.; and Russinovich, M. 2023 · 2023
Cited alongside, same era.
Erasing concepts from diffusion models
Gandikota, R.; Materzynska, J.; Fiotto-Kaufman, J.; and Bau, D. 2023 · 2023
Cited alongside, same era.
A framework for few-shot language model evaluation
Gao, L.; Tow, J.; Abbasi, B.; Biderman, S.; Black, S.; DiPofi, A.; Foster, C.; Golding, L.; Hsu, J.; Le Noac’h, A.; Li, H.; McDonell, K.; Muennighoff, N.; Ociepa, C.; Phang, J.; Reynolds, L.; Schoelkopf, H.; Skowron, A.; Sutawika, L.; Tang, E.; Thite, A.; Wang, B.; Wang, K.; and Zou, A. 2023 · 2023
Cited alongside, same era.
Studying large language model generalization with influence functions
Grosse, R.; Bae, J.; Anil, C.; Elhage, N.; Tamkin, A.; Tajdini, A.; Steiner, B.; Li, D.; Durmus, E.; Perez, E.; et al. 2023 · 2023
Cited alongside, same era.
Knowledge Unlearning for Mitigating Privacy Risks in Language Models
Jang, J.; Yoon, D.; Yang, S.; Cha, S.; Lee, M.; Logeswaran, L.; and Seo, M. 2023 · 2023
Cited alongside, same era.
Jiang, A. Q.; Sablayrolles, A.; Mensch, A.; Bamford, C.; Chaplot, D. S.; Casas, D. d. l.; Bressand, F.; Lengyel, G.; Lample, G.; Saulnier, L.; et al. 2023 · 2023
Cited alongside, same era.
Closest in time.
Fast machine unlearning without retraining through selective synaptic dampening
Foster, J.; Schoepf, S.; and Brintrup, A. 2024 · 2024
Closest in time.
Inexact unlearning needs more careful evaluations to avoid a false sense of privacy
Hayes, J.; Shumailov, I.; Triantafillou, E.; Khalifa, A.; and Papernot, N. 2024 · 2024
Closest in time.
Intrinsic Evaluation of Unlearning Using Parametric Knowledge Traces
Hong, Y.; Yu, L.; Ravfogel, S.; Yang, H.; and Geva, M. 2024 · 2024
Closest in time.
Unlearning Reveals the Influential Training Data of Language Models
Isonuma, M.; and Titov, I. 2024 · 2024
Closest in time.
SoK: Challenges and Opportunities in Federated Unlearning
Jeong, H.; Ma, S.; and Houmansadr, A. 2024 · 2024
Closest in time.
Soul: Unlocking the power of second-order optimization for llm unlearning
Jia, J.; Zhang, Y.; Zhang, Y.; Liu, J.; Runwal, B.; Diffenderfer, J.; Kailkhura, B.; and Liu, S. 2024 · 2024
Closest in time.
Towards Safer Large Language Models through Machine Unlearning
Liu, Z.; Dou, G.; Tan, Z.; Tian, Y.; and Jiang, M. 2024d · 2024
Closest in time.
Eight methods to evaluate robust unlearning in llms
Lynch, A.; Guo, P.; Ewart, A.; Casper, S.; and Hadfield-Menell, D. 2024 · 2024
Closest in time.
Tofu: A task of fictitious unlearning for llms
Maini, P.; Feng, Z.; Schwarzschild, A.; Lipton, Z. C.; and Kolter, J. Z. 2024 · 2024
Closest in time.
Introducing meta llama 3: The most capable openly available llm to date
Meta, A. 2024 · 2024
Closest in time.
Can Sensitive Information Be Deleted From LLMs? Objectives for Defending Against Extraction Attacks
Patil, V.; Hase, P.; and Bansal, M. 2024 · 2024
Closest in time.
In-Context Unlearning: Language Models as Few-Shot Unlearners
Pawelczyk, M.; Neel, S.; and Lakkaraju, H. 2024 · 2024
Closest in time.
Softmax probabilities (mostly) predict large language model correctness on multiple-choice q&a
Plaut, B.; Nguyen, K.; and Trinh, T. 2024 · 2024
Closest in time.
Federated unlearning: A survey on methods, design guidelines, and evaluation metrics
Romandini, N.; Mora, A.; Mazzocca, C.; Montanari, R.; and Bellavista, P. 2024 · 2024
Closest in time.
Unlink to unlearn: Simplifying edge unlearning in gnns
Tan, J.; Sun, F.; Qiu, R.; Su, D.; and Shen, H. 2024 · 2024
Closest in time.
Jailbroken: How does llm safety training fail?
Wei, A.; Haghtalab, N.; and Steinhardt, J. 2024 · 2024
Closest in time.
Yi: Open foundation models by 01. ai
Young, A.; Chen, B.; Li, C.; Huang, C.; Zhang, G.; Zhang, G.; Li, H.; Zhu, J.; Chen, J.; Chang, J.; et al. 2024 · 2024
Closest in time.
GPT-4 Is Too Smart To Be Safe: Stealthy Chat with LLMs via Cipher
Yuan, Y.; Jiao, W.; Wang, W.; tse Huang, J.; He, P.; Shi, S.; and Tu, Z. 2024 · 2024
Closest in time.
Towards efficient and effective unlearning of large language models for recommendation
Wang, H.; Lin, J.; Chen, B.; Yang, Y.; Tang, R.; Zhang, W.; and Yu, Y. 2025 · 2025
Closest in time.