Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are increasingly deployed in real-world applications, raising concerns about the unauthorized use of copyrighted or sensitive data.
Binary codes capable of correcting deletions, insertions, and reversals
V. I. Levenshtein et al · 1966
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
C.-Y. Lin · 2004
Earlier work this paper cites.
Towards making systems forget with machine unlearning
Y. Cao and J. Yang · 2015
Earlier work this paper cites.
Distilling the knowledge in a neural network
G. Hinton · 2015
Earlier work this paper cites.
Semeval-2017 task 1: Semantic textual similarity-multilingual and cross-lingual focused evaluation
D. Cer, M. Diab, E. Agirre, I. Lopez-Gazpio, and L. Specia · 2017
Earlier work this paper cites.
Triviaqa: A large scale distantly supervised challenge dataset for reading comprehension
M. Joshi, E. Choi, D. S. Weld, and L. Zettlemoyer · 2017
Earlier work this paper cites.
The eu general data protection regulation (gdpr)
P. Voigt and A. Von dem Bussche · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord · 2018
Earlier work this paper cites.
Certified data removal from machine learning models
C. Guo, T. Goldstein, A. Hannun, and L. Van Der Maaten · 2019
Earlier work this paper cites.
Eternal sunshine of the spotless net: Selective forgetting in deep networks
A. Golatkar, A. Achille, and S. Soatto · 2020
Earlier work this paper cites.
Measuring massive multitask language understanding
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt · 2020
Earlier work this paper cites.
Machine unlearning for random forests
J. Brophy and D. Lowd · 2021
Earlier work this paper cites.
Extracting training data from large language models
N. Carlini, F. Tramer, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. Brown, D. Song, U. Erlingsson, et al · 2021
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, et al · 2021
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
S. Lin, J. Hilton, and O. Evans · 2021
Earlier work this paper cites.
Bbq: A hand-built bias benchmark for question answering
A. Parrish, A. Chen, N. Nangia, V. Padmakumar, J. Phang, J. Thompson, P. M. Htut, and S. R. Bowman · 2021
Earlier work this paper cites.
4:22-cv-06823, N.D. Cal. 2022
DOE 1 v. GitHub, Inc · 2022
Earlier work this paper cites.
A survey of machine unlearning
T. T. Nguyen, T. T. Huynh, Z. Ren, P. L. Nguyen, A. W.-C. Liew, H. Yin, and Q. V. H. Nguyen · 2022
Earlier work this paper cites.
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al · 2023
Earlier work this paper cites.
Unlearn what you want to forget: Efficient unlearning for llms
J. Chen and D. Yang · 2023
Earlier work this paper cites.
Who’s harry potter? approximate unlearning in llms
R. Eldan and M. Russinovich · 2023
Earlier work this paper cites.
3:23-cv-04625, (N.D. Cal.), 2023
Chabon v. OpenAI, Inc., · 2023
Cited alongside, same era.
3:23-cv-03417, 2023
Kadrey v. Meta Platforms, Inc · 2023
Cited alongside, same era.
23-cv-03416-AMO, (N.D. Cal.), 2023
Tremblay v. OpenAI, Inc., · 2023
Cited alongside, same era.
The times sues openai and microsoft over ai use of copyrighted work
M. M. Grynbaum and R. Mac · 2023
Cited alongside, same era.
Editing models with task arithmetic
G. Ilharco, M. T. Ribeiro, M. Wortsman, S. Gururangan, L. Schmidt, H. Hajishirzi, and A. Farhadi · 2023
Cited alongside, same era.
Knowledge unlearning for mitigating privacy risks in language models
J. Jang, D. Yoon, S. Yang, S. Cha, M. Lee, L. Logeswaran, and M. Seo · 2023
Cited alongside, same era.
Rwku: Benchmarking real-world knowledge unlearning for large language models
Z. Jin, P. Cao, C. Wang, Z. He, H. Yuan, J. Li, Y. Chen, K. Liu, and J. Zhao · 2024
Later among the works it cites.
Towards robust evaluation of unlearning in llms via data transformations
A. Joshi, S. Saha, D. Shukla, S. Vema, H. Jhamtazni, M. Gaur, and A. Modi · 2024
Later among the works it cites.
The wmdp benchmark: Measuring and reducing malicious use with unlearning
N. Li, A. Pan, A. Gopal, S. Yue, D. Berrios, A. Gatti, J. D. Li, A.-K. Dombrowski, S. Goel, L. Phan, et al · 2024
Later among the works it cites.
Large language model unlearning via embedding-corrupted prompts
C. Y. Liu, Y. Wang, J. Flanigan, and Y. Liu · 2024
Later among the works it cites.
Eight methods to evaluate robust unlearning in llms
A. Lynch, P. Guo, A. Ewart, S. Casper, and D. Hadfield-Menell · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Preserving privacy through dememorization: An unlearning technique for mitigating memorization risks in language models
A. M. Kassem, O. A. M. Mahmoud, and S. Saad · 2023
Cited alongside, same era.
Towards unbounded machine unlearning
M. Kurmanji, P. Triantafillou, J. Hayes, and E. Triantafillou · 2023
Cited alongside, same era.
Scalable extraction of training data from (production) language models
M. Nasr, N. Carlini, J. Hayase, M. Jagielski, A. F. Cooper, D. Ippolito, C. A. Choquette-Choo, E. Wallace, F. Tramèr, and K. Lee · 2023
Cited alongside, same era.
Direct preference optimization: Your language model is secretly a reward model
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn · 2023
Cited alongside, same era.
Detecting pretraining data from large language models
W. Shi, A. Ajith, M. Xia, Y. Huang, D. Liu, T. Blevins, D. Chen, and L. Zettlemoyer · 2023
Cited alongside, same era.
To each (textual sequence) its own: Improving memorized-data unlearning in large language models
G.-O. Barbulescu and P. Triantafillou · 2024
Cited alongside, same era.
Tofu: A task of fictitious unlearning for llms
P. Maini, Z. Feng, A. Schwarzschild, Z. C. Lipton, and J. Z. Kolter · 2024
Later among the works it cites.
In-context unlearning: Language models as few shot unlearners
M. Pawelczyk, S. Neel, and H. Lakkaraju · 2024
Later among the works it cites.
Muse: Machine unlearning six-way evaluation for language models
W. Shi, J. Lee, Y. Huang, S. Malladi, J. Zhao, A. Holtzman, D. Liu, L. Zettlemoyer, N. A. Smith, and C. Zhang · 2024
Later among the works it cites.
Position: Llm unlearning benchmarks are weak measures of progress
P. Thaker, S. Hu, N. Kale, Y. Maurya, Z. S. Wu, and V. Smith · 2024
Later among the works it cites.
Guardrail baselines for unlearning in llms
P. Thaker, Y. Maurya, S. Hu, Z. S. Wu, and V. Smith · 2024
Later among the works it cites.
Evaluating copyright takedown methods for language models
B. Wei, W. Shi, Y. Huang, N. A. Smith, C. Zhang, L. Zettlemoyer, K. Li, and P. Henderson · 2024
Later among the works it cites.
Min-k%++: Improved baseline for detecting pre-training data from large language models
J. Zhang, J. Sun, E. Yeats, Y. Ouyang, M. Kuo, J. Zhang, H. F. Yang, and H. Li · 2024
Later among the works it cites.
Negative preference optimization: From catastrophic collapse to effective unlearning
R. Zhang, L. Lin, Y. Bai, and S. Mei · 2024
Later among the works it cites.
Z. Zhang, J. Yang, P. Ke, S. Cui, C. Zheng, H. Wang, and M. Huang · 2024
Later among the works it cites.
On large language model continual unlearning
C. Gao, L. Wang, K. Ding, C. Weng, X. Wang, and Q. Zhu · 2025
Closest in time.
Jogging the memory of unlearned llms through targeted relearning attacks
S. Hu, Y. Fu, S. Wu, and V. Smith · 2025
Closest in time.
Seps: A separability measure for robust unlearning in llms
W. Jeung, S. Yoon, and A. No · 2025
Closest in time.
Rethinking machine unlearning for large language models
S. Liu, Y. Yao, J. Jia, S. Casper, N. Baracaldo, P. Hase, Y. Yao, C. Y. Liu, X. Xu, H. Li, et al · 2025
Closest in time.
Representation bending for large language model safety
A. Yousefpour, T. Kim, R. S. Kwon, S. Lee, W. Jeung, S. Han, A. Wan, H. Ngan, Y. Yu, and J. Choi · 2025
Closest in time.
Catastrophic failure of llm unlearning via quantization
Z. Zhang, F. Wang, X. Li, Z. Wu, X. Tang, H. Liu, Q. He, W. Yin, and S. Wang · 2025
Closest in time.