Fetching the paper…
Reading the bibliography…
Unlearning methods have the potential to improve the privacy and safety of large language models (LLMs) by removing sensitive or harmful information post hoc.
C.-Y. Lin, “Rouge: A package for automatic evaluation of summaries,” in Text summarization branches out , 2004, pp. 74–81
2004
Earlier work this paper cites.
C. Dwork, “Differential privacy,” in International colloquium on automata, languages, and programming . Springer, 2006, pp. 1–12
2006
Earlier work this paper cites.
Y. Cao and J. Yang, “Towards making systems forget with machine unlearning,” in 2015 IEEE symposium on security and privacy . IEEE, 2015, pp. 463–480
2015
Earlier work this paper cites.
2018
Earlier work this paper cites.
A. Ginart, M. Guan, G. Valiant, and J. Y. Zou, “Making ai forget you: Data deletion in machine learning,” Advances in neural information processing systems , vol. 32, 2019
2019
Earlier work this paper cites.
S. Shintre, K. A. Roundy, and J. Dhaliwal, “Making machine learning forget,” in Privacy Technologies and Policy: 7th Annual Privacy Forum, APF 2019, Rome, Italy, June 13–14, 2019, Proceedings 7 . Springer, 2019, pp. 72–83
2019
Earlier work this paper cites.
2019
Earlier work this paper cites.
2020
Earlier work this paper cites.
A. Golatkar, A. Achille, and S. Soatto, “Eternal sunshine of the spotless net: Selective forgetting in deep networks,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2020, pp. 9304–9312
2020
Earlier work this paper cites.
C. Guo, T. Goldstein, A. Hannun, and L. Van Der Maaten, “Certified data removal from machine learning models,” in International Conference on Machine Learning , 2020
2020
Earlier work this paper cites.
S. Garg, S. Goldwasser, and P. N. Vasudevan, “Formalizing data deletion in the context of the right to be forgotten,” in Annual International Conference on the Theory and Applications of Cryptographic Techniques . Springer, 2020, pp. 373–402
2020
Earlier work this paper cites.
L. Bourtoule, V. Chandrasekaran, C. A. Choquette-Choo, H. Jia, A. Travers, B. Zhang, D. Lie, and N. Papernot, “Machine unlearning,” in 2021 IEEE Symposium on Security and Privacy (SP) . IEEE, 2021, pp. 141–159
2021
Earlier work this paper cites.
V. Gupta, C. Jung, S. Neel, A. Roth, S. Sharifi-Malvajerdi, and C. Waites, “Adaptive machine unlearning,” Advances in Neural Information Processing Systems , vol. 34, pp. 16 319–16 330, 2021
2021
Earlier work this paper cites.
Z. Zhao, E. Wallace, S. Feng, D. Klein, and S. Singh, “Calibrate before use: Improving few-shot performance of language models,” in International conference on machine learning . PMLR, 2021, pp. 12 697–12 706
2021
Earlier work this paper cites.
S. Neel, A. Roth, and S. Sharifi-Malvajerdi, “Descent-to-delete: Gradient-based methods for machine unlearning,” in Algorithmic Learning Theory . PMLR, 2021, pp. 931–962
2021
Earlier work this paper cites.
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt, “Measuring massive multitask language understanding,” in International Conference on Learning Representations , 2021
2021
Earlier work this paper cites.
2021
Earlier work this paper cites.
A. Sekhari, J. Acharya, G. Kamath, and A. T. Suresh, “Remember what you want to forget: Algorithms for machine unlearning,” Advances in Neural Information Processing Systems , vol. 34, pp. 18 075–18 086, 2021
2021
Earlier work this paper cites.
H. Brown, K. Lee, F. Mireshghallah, R. Shokri, and F. Tramèr, “What does it mean for a language model to preserve privacy?” in Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency , 2022, pp. 2280–2292
2022
Earlier work this paper cites.
A. Thudi, G. Deza, V. Chandrasekaran, and N. Papernot, “Unrolling sgd: Understanding factors influencing machine unlearning,” in 2022 IEEE 7th European Symposium on Security and Privacy (EuroS&P) . IEEE, 2022, pp. 303–319
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
M. Chen, Z. Zhang, T. Wang, M. Backes, M. Humbert, and Y. Zhang, “Graph unlearning,” in Proceedings of the 2022 ACM SIGSAC conference on computer and communications security , 2022, pp. 499–513
2022
Earlier work this paper cites.
F. Tramèr, G. Kamath, and N. Carlini, “Position: Considerations for differentially private learning with large-scale public pretraining,” in Forty-first International Conference on Machine Learning , 2022
2022
Earlier work this paper cites.
N. Carlini, M. Jagielski, C. Zhang, N. Papernot, A. Terzis, and F. Tramer, “The privacy onion effect: Memorization is relative,” Advances in Neural Information Processing Systems , vol. 35, pp. 13 263–13 276, 2022
2022
Earlier work this paper cites.
2022
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
2023
Earlier work this paper cites.
H. Xu, T. Zhu, L. Zhang, W. Zhou, and P. S. Yu, “Machine unlearning: A survey,” ACM Computing Surveys , vol. 56, no. 1, pp. 1–36, 2023
2023
Earlier work this paper cites.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
N. Carlini, H. Cooper, F. Tramer, and C. Zhang, “Lm extraction benchmark,” https://github.com/google-research/lm-extraction-benchmark , 2023, accessed: 2024-09-24
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
P. Hase, M. Bansal, B. Kim, and A. Ghandeharioun, “Does localization inform editing? surprising differences in causality-based localization vs,” Knowledge Editing in Language Models , 2023
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2023
Cited alongside, same era.
A. Cohen, A. Smith, M. Swanberg, and P. N. Vasudevan, “Control, confidentiality, and the right to be forgotten,” in Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security , 2023, pp. 3358–3372
2023
Cited alongside, same era.
2023
Cited alongside, same era.
2024
Cited alongside, same era.
N. Li, A. Pan, A. Gopal, S. Yue, D. Berrios, A. Gatti, J. D. Li, A.-K. Dombrowski, S. Goel, L. Phan et al. , “The WMDP benchmark: Measuring and reducing malicious use with unlearning,” in International Conference on Machine Learning , 2024
2024
Cited alongside, same era.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
R. Tamirisa, B. Bharathi, A. Zhou, and B. L. M. Mazeika, “Toward robust unlearning for llms,” in ICLR 2024 Workshop on Secure and Trustworthy Large Language Models , 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn, “Direct preference optimization: Your language model is secretly a reward model,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.
2024
Closest in time.