Fetching the paper…
Reading the bibliography…
We explore machine unlearning (MU) in the domain of large language models (LLMs), referred to as LLM unlearning.
Towards making systems forget with machine unlearning
Cao, Y. and Yang, J · 2015
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Understanding black-box predictions via influence functions
Koh, P. W. and Liang, P · 2017
Earlier work this paper cites.
Towards deep learning models resistant to adversarial attacks
Madry, A., Makelov, A., Schmidt, L., Tsipras, D., and Vladu, A · 2017
Earlier work this paper cites.
Membership inference attacks against machine learning models
Shokri, R., Stronati, M., Song, C., and Shmatikov, V · 2017
Earlier work this paper cites.
Obfuscated gradients give a false sense of security: Circumventing defenses to adversarial examples
Athalye, A., Carlini, N., and Wagner, D · 2018
Earlier work this paper cites.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Carlini, N., Liu, C., Erlingsson, Ú., Kos, J., and Song, D · 2019
Earlier work this paper cites.
Making ai forget you: Data deletion in machine learning
Ginart, A., Guan, M., Valiant, G., and Zou, J. Y · 2019
Earlier work this paper cites.
Certified data removal from machine learning models
Guo, C., Goldstein, T., Hannun, A., and Van Der Maaten, L · 2019
Earlier work this paper cites.
Unlearn dataset bias in natural language inference by fitting the residual
He, H., Zha, S., and Wang, H · 2019
Earlier work this paper cites.
The european union general data protection regulation: what it is and what it means
Hoofnagle, C. J., van der Sloot, B., and Borgesius, F. Z · 2019
Earlier work this paper cites.
Harnessing the vulnerability of latent layers in adversarially trained models
Kumari, N., Singh, M., Sinha, A., Machiraju, H., Krishnamurthy, B., and Balasubramanian, V. N · 2019
Earlier work this paper cites.
Adversarial training for free!
Shafahi, A., Najibi, M., Ghiasi, M. A., Xu, Z., Dickerson, J., Studer, C., Davis, L. S., Taylor, G., and Goldstein, T · 2019
Earlier work this paper cites.
Neural cleanse: Identifying and mitigating backdoor attacks in neural networks
Wang, B., Yao, Y., Shan, S., Li, H., Viswanath, B., Zheng, H., and Zhao, B. Y · 2019
Earlier work this paper cites.
Freelb: Enhanced adversarial training for natural language understanding
Zhu, C., Cheng, Y., Gan, Z., Sun, S., Goldstein, T., and Liu, J · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Realtoxicityprompts: Evaluating neural toxic degeneration in language models
Gehman, S., Gururangan, S., Sap, M., Choi, Y., and Smith, N. A · 2020
Earlier work this paper cites.
Transformer feed-forward layers are key-value memories
Geva, M., Schuster, R., Berant, J., and Levy, O · 2020
Earlier work this paper cites.
Eternal sunshine of the spotless net: Selective forgetting in deep networks
Golatkar, A., Achille, A., and Soatto, S · 2020
Earlier work this paper cites.
Liu, G., Ma, X., Yang, Y., Wang, C., and Liu, J · 2020
Earlier work this paper cites.
Fast is better than free: Revisiting adversarial training
Wong, E., Rice, L., and Kolter, J. Z · 2020
Earlier work this paper cites.
On the dangers of stochastic parrots: Can language models be too big?
Bender, E. M., Gebru, T., McMillan-Major, A., and Shmitchell, S · 2021
Earlier work this paper cites.
Machine unlearning
Bourtoule, L., Chandrasekaran, V., Choquette-Choo, C. A., Jia, H., Travers, A., Zhang, B., Lie, D., and Papernot, N · 2021
Earlier work this paper cites.
Extracting training data from large language models
Carlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, U., et al · 2021
Earlier work this paper cites.
Knowledge neurons in pretrained transformers
Dai, D., Dong, L., Hao, Y., Sui, Z., Chang, B., and Wei, F · 2021
Earlier work this paper cites.
Approximate data deletion from machine learning models
Izzo, Z., Smart, M. A., Chaudhuri, K., and Zou, J · 2021
Earlier work this paper cites.
Anti-backdoor learning: Training clean models on poisoned data
Li, Y., Lyu, X., Koren, N., Lyu, L., Li, B., and Ma, X · 2021
Earlier work this paper cites.
Descent-to-delete: Gradient-based methods for machine unlearning
Neel, S., Roth, A., and Sharifi-Malvajerdi, S · 2021
Earlier work this paper cites.
Remember what you want to forget: Algorithms for machine unlearning
Sekhari, A., Acharya, J., Kamath, G., and Suresh, A. T · 2021
Earlier work this paper cites.
Machine unlearning via algorithmic stability
Ullah, E., Mai, T., Rao, A., Rossi, R. A., and Arora, R · 2021
Earlier work this paper cites.
Machine unlearning of features and labels
Warnecke, A., Pirch, L., Wressnegger, C., and Rieck, K · 2021
Earlier work this paper cites.
If influence functions are the answer, then what is the question?
Bae, J., Ng, N., Lo, A., Ghassemi, M., and Grosse, R. B · 2022
Earlier work this paper cites.
Constitutional ai: Harmlessness from ai feedback
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., McKinnon, C., et al · 2022
Earlier work this paper cites.
Evaluating machine unlearning via epistemic uncertainty
Becker, A. and Liebig, T · 2022
Earlier work this paper cites.
Quantifying memorization across neural language models
Carlini, N., Ippolito, D., Jagielski, M., Lee, K., Tramer, F., and Zhang, C · 2022
Earlier work this paper cites.
Recommendation unlearning
Chen, C., Sun, F., Zhang, M., and Ding, B · 2022
Earlier work this paper cites.
Graph unlearning
Chen, M., Zhang, Z., Wang, T., Backes, M., Humbert, M., and Zhang, Y · 2022
Earlier work this paper cites.
Efficient model updates for approximate unlearning of graph-structured data
Chien, E., Pan, C., and Milenkovic, O · 2022
Earlier work this paper cites.
Federated unlearning: How to efficiently erase a client in fl?
Halimi, A., Kadhe, S., Rawat, A., and Baracaldo, N · 2022
Earlier work this paper cites.
Editing models with task arithmetic
Ilharco, G., Ribeiro, M. T., Wortsman, M., Gururangan, S., Schmidt, L., Hajishirzi, H., and Farhadi, A · 2022
Earlier work this paper cites.
Knowledge unlearning for mitigating privacy risks in language models
Jang, J., Yoon, D., Yang, S., Cha, S., Lee, M., Logeswaran, L., and Seo, M · 2022
Earlier work this paper cites.
Privacy adhering machine un-learning in nlp
Kumar, V. B., Gangadharaiah, R., and Roth, D · 2022
Earlier work this paper cites.
Large language models with controllable working memory
Li, D., Rawat, A. S., Zaheer, M., Wang, X., Lukasik, M., Veit, A., Yu, F., and Kumar, S · 2022
Earlier work this paper cites.
Algorithmic destruction
Li, T. C · 2022
Earlier work this paper cites.
Backdoor defense with machine unlearning
Liu, Y., Fan, M., Chen, C., Liu, X., Ma, Z., Wang, L., and Ma, J · 2022
Earlier work this paper cites.
Quark: Controllable text generation with reinforced unlearning
Lu, X., Welleck, S., Hessel, J., Jiang, L., Qin, L., West, P., Ammanabrolu, P., and Choi, Y · 2022
Cited alongside, same era.
Memory-assisted prompt editing to improve gpt-3 after deployment
Madaan, A., Tandon, N., Clark, P., and Yang, Y · 2022
Cited alongside, same era.
Locating and editing factual associations in gpt
Meng, K., Bau, D., Andonian, A., and Belinkov, Y · 2022
Cited alongside, same era.
Memory-based model editing at scale
Mitchell, E., Lin, C., Bosselut, A., Manning, C. D., and Finn, C · 2022
Cited alongside, same era.
A survey of machine unlearning
Nguyen, T. T., Huynh, T. T., Nguyen, P. L., Liew, A. W.-C., Yin, H., and Nguyen, Q. V. H · 2022
Cited alongside, same era.
Fine-tuning aligned language models compromises safety, even when users do not intend to!
Qi, X., Zeng, Y., Xie, T., Chen, P.-Y., Jia, R., Mittal, P., and Henderson, P · 2023
Later among the works it cites.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., and Finn, C · 2023
Later among the works it cites.
Universal jailbreak backdoors from poisoned human feedback
Rando, J. and Tramèr, F · 2023
Later among the works it cites.
Adversarial training should be cast as a non-zero-sum game
Robey, A., Latorre, F., Pappas, G. J., Hassani, H., and Cevher, V · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
Discovering language model behaviors with model-written evaluations
Perez, E., Ringer, S., Lukošiūtė, K., Nguyen, K., Chen, E., Heiner, S., Pettit, C., Olsson, C., Kundu, S., Kadavath, S., et al · 2022
Cited alongside, same era.
Fair infinitesimal jackknife: Mitigating the influence of biased training data points without refitting
Sattigeri, P., Ghosh, S., Padhi, I., Dognin, P., and Varshney, K. R · 2022
Cited alongside, same era.
Unrolling sgd: Understanding factors influencing machine unlearning
Thudi, A., Deza, G., Chandrasekaran, V., and Papernot, N · 2022
Cited alongside, same era.
Federated unlearning via class-discriminative pruning
Wang, J., Guo, S., Xie, X., and Qi, H · 2022
Cited alongside, same era.
Emergent abilities of large language models
Wei, J., Tay, Y., Bommasani, R., Raffel, C., Zoph, B., Borgeaud, S., Yogatama, D., Bosma, M., Zhou, D., Metzler, D., et al · 2022
Cited alongside, same era.
Revisiting and advancing fast adversarial training through the lens of bi-level optimization
Zhang, Y., Zhang, G., Khanduri, P., Hong, M., Chang, S., and Liu, S · 2022
Cited alongside, same era.
Shevlane, T., Farquhar, S., Garfinkel, B., Phuong, M., Whittlestone, J., Leung, J., Kokotajlo, D., Marchal, N., Anderljung, M., Kolt, N., et al · 2023
Later among the works it cites.
Detecting pretraining data from large language models
Shi, W., Ajith, A., Xia, M., Huang, Y., Liu, D., Blevins, T., Chen, D., and Zettlemoyer, L · 2023
Later among the works it cites.
Knowledge unlearning for llms: Tasks, methods, and challenges
Si, N., Zhang, H., Chang, H., Zhang, W., Qu, D., and Zhang, W · 2023
Later among the works it cites.
Evaluating and mitigating discrimination in language model decisions
Tamkin, A., Askell, A., Lovitt, L., Durmus, E., Joseph, N., Kravec, S., Nguyen, K., Kaplan, J., and Ganguli, D · 2023
Later among the works it cites.
Kga: A general machine unlearning framework based on knowledge gap alignment
Wang, L., Chen, T., Yuan, W., Zeng, X., Wong, K.-F., and Yin, H · 2023
Later among the works it cites.
Jailbroken: How does llm safety training fail?
Wei, A., Haghtalab, N., and Steinhardt, J · 2023
Later among the works it cites.
Unveiling the implicit toxicity in large language models
Wen, J., Ke, P., Sun, H., Zhang, Z., Li, C., Bai, J., and Huang, M · 2023
Later among the works it cites.
Netflix and forget: Efficient and exact machine unlearning from bi-linear recommendations
Xu, M., Sun, J., Yang, X., Yao, K., and Wang, C · 2023
Later among the works it cites.
Shadow alignment: The ease of subverting safely-aligned language models
Yang, X., Wang, X., Zhang, Q., Petzold, L., Wang, W. Y., Zhao, X., and Lin, D · 2023
Later among the works it cites.
Large language model unlearning
Yao, Y., Xu, X., and Liu, Y · 2023
Later among the works it cites.
Low-resource languages jailbreak gpt-4, 2023
Yong, Z.-X., Menghini, C., and Bach, S. H · 2023
Later among the works it cites.
Unlearning bias in language models by partitioning gradients
Yu, C., Jeoung, S., Kasi, A., Yu, P., and Ji, H · 2023
Later among the works it cites.
Rrhf: Rank responses to align language models with human feedback without tears
Yuan, Z., Yuan, H., Tan, C., Wang, W., Huang, S., and Huang, F · 2023
Later among the works it cites.
Can we edit factual knowledge by in-context learning?
Zheng, C., Li, L., Dong, Q., Fan, Y., Wu, Z., Xu, J., and Chang, B · 2023
Later among the works it cites.
Mquake: Assessing knowledge editing in language models via multi-hop questions
Zhong, Z., Wu, Z., Manning, C. D., Potts, C., and Chen, D · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M · 2023
Later among the works it cites.
To each (textual sequence) its own: Improving memorized-data unlearning in large language models
Barbulescu, G.-O. and Triantafillou, P · 2024
Closest in time.
From hope to safety: Unlearning biases of deep models via gradient penalization in latent space
Dreyer, M., Pahde, F., Anders, C. J., Samek, W., and Lapuschkin, S · 2024
Closest in time.
Do membership inference attacks work on large language models?
Duan, M., Suri, A., Mireshghallah, N., Min, S., Shi, W., Zettlemoyer, L., Tsvetkov, Y., Choi, Y., Evans, D., and Hajishirzi, H · 2024
Closest in time.
The times sues openai and microsoft over a.i. use of copyrighted work
Grynbaum, M. and Mac, R · 2024
Closest in time.
Jogging the memory of unlearned model through targeted relearning attack
Hu, S., Fu, Y., Wu, Z. S., and Smith, V · 2024
Closest in time.
Sleeper agents: Training deceptive llms that persist through safety training
Hubinger, E., Denison, C., Mu, J., Lambert, M., Tong, M., MacDiarmid, M., Lanham, T., Ziegler, D. M., Maxwell, T., Cheng, N., et al · 2024
Closest in time.
SOUL: Unlocking the power of second-order optimization for LLM unlearning
Jia, J., Zhang, Y., Zhang, Y., Liu, J., Runwal, B., Diffenderfer, J., Kailkhura, B., and Liu, S · 2024
Closest in time.
Protecting privacy through approximating optimal parameters for sequence unlearning in language models
Lee, D., Rim, D., Choi, M., and Choo, J · 2024
Closest in time.
Large language model unlearning via embedding-corrupted prompts
Liu, C. Y., Wang, Y., Flanigan, J., and Liu, Y · 2024
Closest in time.
An adversarial perspective on machine unlearning for ai safety
Łucki, J., Wei, B., Huang, Y., Henderson, P., Tramèr, F., and Rando, J · 2024
Closest in time.
Eight methods to evaluate robust unlearning in llms
Lynch, A., Guo, P., Ewart, A., Casper, S., and Hadfield-Menell, D · 2024
Closest in time.
TOFU: A task of fictitious unlearning for LLMs
Maini, P., Feng, Z., Schwarzschild, A., Lipton, Z. C., and Kolter, J. Z · 2024
Closest in time.
Unlearnable algorithms for in-context learning
Muresanu, A., Thudi, A., Zhang, M. R., and Papernot, N · 2024
Closest in time.
Can sensitive information be deleted from llms? objectives for defending against extraction attacks
Patil, V., Hase, P., and Bansal, M · 2024
Closest in time.
Machine unlearning for recommendation systems: An insight
Sachdeva, B., Rathee, H., Sharma, A., Wydmański, W., et al · 2024
Closest in time.
Are emergent abilities of large language models a mirage?
Schaeffer, R., Miranda, B., and Koyejo, S · 2024
Closest in time.
Muse: Machine unlearning six-way evaluation for language models
Shi, W., Lee, J., Huang, Y., Malladi, S., Zhao, J., Holtzman, A., Liu, D., Zettlemoyer, L., Smith, N. A., and Zhang, C · 2024
Closest in time.
Ununlearning: Unlearning is not sufficient for content regulation in advanced generative ai
Shumailov, I., Hayes, J., Triantafillou, E., Ortiz-Jimenez, G., Papernot, N., Jagielski, M., Yona, I., Howard, H., and Bagdasaryan, E · 2024
Closest in time.
Sarah silverman sues openai and meta over copyright infringement
Small, Z · 2024
Closest in time.
Data forging is harder than you think
Suliman, M., Kadhe, S., Halimi, A., Leith, D., Baracaldo, N., and Rawat, A · 2024
Closest in time.
Tamper-resistant safeguards for open-weight llms
Tamirisa, R., Bharathi, B., Phan, L., Zhou, A., Gatti, A., Suresh, T., Lin, M., Wang, J., Wang, R., Arel, R., et al · 2024
Closest in time.
Guardrail baselines for unlearning in llms
Thaker, P., Maurya, Y., and Smith, V · 2024
Closest in time.
Machine unlearning of pre-trained large language models
Yao, J., Chien, E., Du, M., Niu, X., Wang, T., Cheng, Z., and Yue, X · 2024
Closest in time.
What makes unlearning hard and what to do about it
Zhao, K., Kurmanji, M., Bărbulescu, G.-O., Triantafillou, E., and Triantafillou, P · 2024
Closest in time.