Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) embed sensitive, human-generated data, prompting the need for unlearning methods.
Introducing the enron corpus
Klimt, B. and Yang, Y · 2004
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
Towards making systems forget with machine unlearning
Cao, Y. and Yang, J · 2015
Earlier work this paper cites.
Deep learning with differential privacy
Abadi, M., Chu, A., Goodfellow, I., McMahan, H. B., Mironov, I., Talwar, K., and Zhang, L · 2016
Earlier work this paper cites.
Membership inference attacks against machine learning models
Shokri, R., Stronati, M., Song, C., and Shmatikov, V · 2017
Earlier work this paper cites.
GDPR: General Data Protection Regulation (EU) 2016/679: Post-reform Personal Data Protection in the European Union
Krzysztofek, M · 2018
Earlier work this paper cites.
Privacy risk in machine learning: Analyzing the connection to overfitting
Yeom, S., Giacomelli, I., Fredrikson, M., and Jha, S · 2018
Earlier work this paper cites.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Carlini, N., Liu, C., Erlingsson, Ú., Kos, J., and Song, D · 2019
Earlier work this paper cites.
Neural legal judgment prediction in english
Chalkidis, I., Androutsopoulos, I., and Aletras, N · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y · 2019
Earlier work this paper cites.
What neural networks memorize and why: Discovering the long tail via influence estimation
Feldman, V. and Zhang, C · 2020
Earlier work this paper cites.
Forgetting outside the box: Scrubbing deep networks of information accessible from input-output observations
Golatkar, A., Achille, A., and Soatto, S · 2020
Earlier work this paper cites.
Certified data removal from machine learning models
Guo, C., Goldstein, T., Hannun, A., and Van Der Maaten, L · 2020
Earlier work this paper cites.
Auditing differentially private machine learning: How private is private sgd?
Jagielski, M., Ullman, J., and Oprea, A · 2020
Earlier work this paper cites.
The pitfalls of average-case differential privacy
Steinke, T. and Ullman, J · 2020
Earlier work this paper cites.
Machine unlearning
Bourtoule, L., Chandrasekaran, V., Choquette-Choo, C. A., Jia, H., Travers, A., Zhang, B., Lie, D., and Papernot, N · 2021
Earlier work this paper cites.
Extracting training data from large language models
Carlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, U., et al · 2021
Earlier work this paper cites.
Amnesiac machine learning
Graves, L., Nagisetty, V., and Ganesh, V · 2021
Cited alongside, same era.
Lora: Low-rank adaptation of large language models
Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., and Chen, W · 2021
Cited alongside, same era.
Adversary instantiation: Lower bounds for differentially private machine learning
Nasr, M., Songi, S., Thakurta, A., Papernot, N., and Carlin, N · 2021
Cited alongside, same era.
Descent-to-delete: Gradient-based methods for machine unlearning
Neel, S., Roth, A., and Sharifi-Malvajerdi, S · 2021
Cited alongside, same era.
Remember what you want to forget: Algorithms for machine unlearning
Sekhari, A., Acharya, J., Kamath, G., and Suresh, A. T · 2021
Cited alongside, same era.
Machine unlearning via algorithmic stability
Tight auditing of differentially private machine learning
Nasr, M., Hayes, J., Steinke, T., Balle, B., Tramèr, F., Jagielski, M., Carlini, N., and Terzis, A · 2023
Later among the works it cites.
Detecting pretraining data from large language models
Shi, W., Ajith, A., Xia, M., Huang, Y., Liu, D., Blevins, T., Chen, D., and Zettlemoyer, L · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
From adaptive query release to machine unlearning
Ullah, E. and Arora, R · 2023
Later among the works it cites.
Depn: Detecting and editing privacy neurons in pretrained language models
Wu, X., Li, J., Xu, M., Dong, W., Wu, S., Bian, C., and Xiong, D · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ullah, E., Mai, T., Rao, A., Rossi, R. A., and Arora, R · 2021
Cited alongside, same era.
Membership inference attacks from first principles
Carlini, N., Chien, S., Nasr, M., Song, S., Terzis, A., and Tramer, F · 2022
Cited alongside, same era.
Towards adversarial evaluations for inexact machine unlearning
Goel, S., Prabhu, A., Sanyal, A., Lim, S.-N., Torr, P., and Kumaraguru, P · 2022
Cited alongside, same era.
Knowledge unlearning for mitigating privacy risks in language models
Jang, J., Yoon, D., Yang, S., Cha, S., Lee, M., Logeswaran, L., and Seo, M · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al · 2022
Cited alongside, same era.
On the necessity of auditable algorithmic definitions for machine unlearning
Thudi, A., Jia, H., Shumailov, I., and Papernot, N · 2022
Cited alongside, same era.
Enhanced membership inference attacks against machine learning models
Ye, J., Maddi, A., Murakonda, S. K., Bindschaedler, V., and Shokri, R · 2022
Cited alongside, same era.
Later among the works it cites.
Large language model unlearning
Yao, Y., Xu, X., and Liu, Y · 2023
Later among the works it cites.
Evaluations of machine learning privacy defenses are misleading
Aerni, M., Zhang, J., and Tramèr, F · 2024
Closest in time.
Do membership inference attacks work on large language models?
Duan, M., Suri, A., Mireshghallah, N., Min, S., Shi, W., Zettlemoyer, L., Tsvetkov, Y., Choi, Y., Evans, D., and Hajishirzi, H · 2024
Closest in time.
Inexact unlearning needs more careful evaluations to avoid a false sense of privacy
Hayes, J., Shumailov, I., Triantafillou, E., Khalifa, A., and Papernot, N · 2024
Closest in time.
Towards unbounded machine unlearning
Kurmanji, M., Triantafillou, P., Hayes, J., and Triantafillou, E · 2024
Closest in time.
Breaking the trilemma of privacy, utility, and efficiency via controllable machine unlearning
Liu, Z., Dou, G., Chien, E., Zhang, C., Tian, Y., and Zhu, Z · 2024
Closest in time.
Tofu: A task of fictitious unlearning for llms
Maini, P., Feng, Z., Schwarzschild, A., Lipton, Z. C., and Kolter, J. Z · 2024
Closest in time.
Privacy auditing with one (1) training run
Steinke, T., Nasr, M., and Jagielski, M · 2024
Closest in time.
Gradients look alike: Sensitivity is often overestimated in { \{ DP-SGD } \}
Thudi, A., Jia, H., Meehan, C., Shumailov, I., and Papernot, N · 2024
Closest in time.
Machine unlearning of pre-trained large language models
Yao, J., Chien, E., Du, M., Niu, X., Wang, T., Cheng, Z., and Yue, X · 2024
Closest in time.
Negative preference optimization: From catastrophic collapse to effective unlearning
Zhang, R., Lin, L., Bai, Y., and Mei, S · 2024
Closest in time.
What makes unlearning hard and what to do about it
Zhao, K., Kurmanji, M., Bărbulescu, G.-O., Triantafillou, E., and Triantafillou, P · 2024
Closest in time.