Fetching the paper…
Reading the bibliography…
Language Models (LMs) are prone to ''memorizing'' training data, including substantial sensitive user information.
Rouge: A package for automatic evaluation of summaries
Lin, C.-Y · 2004
Earlier work this paper cites.
The right to be forgotten
Rosen, J · 2011
Earlier work this paper cites.
Adam: A method for stochastic optimization
Kingma, D. P. and Ba, J · 2014
Earlier work this paper cites.
Lightgbm: A highly efficient gradient boosting decision tree
Ke, G., Meng, Q., Finley, T., Wang, T., Chen, W., Ma, W., Ye, Q., and Liu, T.-Y · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
Loshchilov, I · 2017
Earlier work this paper cites.
Membership inference attacks against machine learning models
Shokri, R., Stronati, M., Song, C., and Shmatikov, V · 2017
Earlier work this paper cites.
General data protection regulation
GDPR, G. D. P. R · 2018
Earlier work this paper cites.
Amazon reportedly employs thousands of people to listen to your alexa conversations, 2019
Valinsky, J · 2019
Earlier work this paper cites.
Machine unlearning
Bourtoule, L., Chandrasekaran, V., Choquette-Choo, C. A., Jia, H., Travers, A., Zhang, B., Lie, D., and Papernot, N · 2021
Earlier work this paper cites.
Extracting training data from large language models
Carlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, U., et al · 2021
Earlier work this paper cites.
When machine unlearning jeopardizes privacy
Chen, M., Zhang, Z., Wang, T., Backes, M., Humbert, M., and Zhang, Y · 2021
Earlier work this paper cites.
Lamp: Extracting text from gradients with language model priors
Balunovic, M., Dimitrov, D., Jovanović, N., and Vechev, M · 2022
Earlier work this paper cites.
Membership inference attacks from first principles
Carlini, N., Chien, S., Nasr, M., Song, S., Terzis, A., and Tramer, F · 2022
Cited alongside, same era.
Editing models with task arithmetic
Ilharco, G., Ribeiro, M. T., Wortsman, M., Gururangan, S., Schmidt, L., Hajishirzi, H., and Farhadi, A · 2022
Cited alongside, same era.
Knowledge unlearning for mitigating privacy risks in language models
Jang, J., Yoon, D., Yang, S., Cha, S., Lee, M., Logeswaran, L., and Seo, M · 2022
Cited alongside, same era.
Pythia: A suite for analyzing large language models across training and scaling
Biderman, S., Schoelkopf, H., Anthony, Q. G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., et al · 2023
Cited alongside, same era.
Unlearn what you want to forget: Efficient unlearning for llms
Inexact unlearning needs more careful evaluations to avoid a false sense of privacy
Hayes, J., Shumailov, I., Triantafillou, E., Khalifa, A., and Papernot, N · 2024
Closest in time.
Learn what you want to unlearn: Unlearning inversion attacks against machine unlearning
Hu, H., Wang, S., Dong, T., and Xue, M · 2024
Closest in time.
Unlearn and burn: Adversarial machine unlearning requests destroy model accuracy
Huang, Y., Liu, D., Chua, L., Ghazi, B., Kamath, P., Kumar, R., Manurangsi, P., Nasr, M., Sinha, A., and Zhang, C · 2024
Closest in time.
Tofu: A task of fictitious unlearning for llms
Maini, P., Feng, Z., Schwarzschild, A., Lipton, Z. C., and Kolter, J. Z · 2024
Closest in time.
Can gpt-4o be trusted with your private data?, 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chen, J. and Yang, D · 2023
Cited alongside, same era.
Trustworthy llms: a survey and guideline for evaluating large language models’ alignment
Liu, Y., Yao, Y., Ton, J.-F., Zhang, X., Cheng, R. G. H., Klochkov, Y., Taufiq, M. F., and Li, H · 2023
Cited alongside, same era.
Membership inference attacks against language models via neighbourhood comparison
Mattern, J., Mireshghallah, F., Jin, Z., Schölkopf, B., Sachan, M., and Berg-Kirkpatrick, T · 2023
Cited alongside, same era.
Chatgpt (mar 14 version) [large language model], 2023
OpenAI · 2023
Cited alongside, same era.
In-context unlearning: Language models as few shot unlearners
Pawelczyk, M., Neel, S., and Lakkaraju, H · 2023
Cited alongside, same era.
Evaluations of machine learning privacy defenses are misleading
Aerni, M., Zhang, J., and Tramèr, F · 2024
Cited alongside, same era.
Do membership inference attacks work on large language models?
Duan, M., Suri, A., Mireshghallah, N., Min, S., Shi, W., Zettlemoyer, L., Tsvetkov, Y., Choi, Y., Evans, D., and Hajishirzi, H · 2024
Cited alongside, same era.
Membership inference attacks cannot prove that a model was trained on your data
Zhang, J., Das, D., Kamath, G., and Tramèr, F
Cited in the paper.
O’Flaherty, K · 2024
Closest in time.
Machine unlearning fails to remove data poisoning attacks
Pawelczyk, M., Di, J. Z., Lu, Y., Kamath, G., Sekhari, A., and Neel, S · 2024
Closest in time.
Direct preference optimization: Your language model is secretly a reward model
Rafailov, R., Sharma, A., Mitchell, E., Manning, C. D., Ermon, S., and Finn, C · 2024
Closest in time.
Muse: Machine unlearning six-way evaluation for language models
Shi, W., Lee, J., Huang, Y., Malladi, S., Zhao, J., Holtzman, A., Liu, D., Zettlemoyer, L., Smith, N. A., and Zhang, C · 2024
Closest in time.
Ununlearning: Unlearning is not sufficient for content regulation in advanced generative ai
Shumailov, I., Hayes, J., Triantafillou, E., Ortiz-Jimenez, G., Papernot, N., Jagielski, M., Yona, I., Howard, H., and Bagdasaryan, E · 2024
Closest in time.
A synthetic dataset for personal attribute inference
Yukhymenko, H., Staab, R., Vero, M., and Vechev, M · 2024
Closest in time.
Opt: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., et al · 2024
Closest in time.