Fetching the paper…
Reading the bibliography…
We study how to perform unlearning, i.e.
Bleu: a method for automatic evaluation of machine translation
K. Papineni, S. Roukos, T. Ward, and W.-J. Zhu · 2002
Earlier work this paper cites.
Towards making systems forget with machine unlearning
Y. Cao and J. Yang · 2015
Earlier work this paper cites.
Aligning books and movies: Towards story-like visual explanations by watching movies and reading books
Y. Zhu, R. Kiros, R. Zemel, R. Salakhutdinov, R. Urtasun, A. Torralba, and S. Fidler · 2015
Earlier work this paper cites.
Deep reinforcement learning from human preferences
P. F. Christiano, J. Leike, T. Brown, M. Martic, S. Legg, and D. Amodei · 2017
Earlier work this paper cites.
Proximal policy optimization algorithms
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov · 2017
Earlier work this paper cites.
Membership inference attacks against machine learning models
R. Shokri, M. Stronati, C. Song, and V. Shmatikov · 2017
Earlier work this paper cites.
Certified data removal from machine learning models
C. Guo, T. Goldstein, A. Hannun, and L. Van Der Maaten · 2019
Earlier work this paper cites.
The curious case of neural text degeneration
A. Holtzman, J. Buys, L. Du, M. Forbes, and Y. Choi · 2019
Earlier work this paper cites.
When does label smoothing help?
R. Müller, S. Kornblith, and G. E. Hinton · 2019
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
T. Zhang, V. Kishore, F. Wu, K. Q. Weinberger, and Y. Artzi · 2019
Earlier work this paper cites.
Bleurt: Learning robust metrics for text generation
T. Sellam, D. Das, and A. P. Parikh · 2020
Earlier work this paper cites.
Machine unlearning
L. Bourtoule, V. Chandrasekaran, C. A. Choquette-Choo, H. Jia, A. Travers, B. Zhang, D. Lie, and N. Papernot · 2021
Earlier work this paper cites.
Extracting training data from large language models
N. Carlini, F. Tramer, E. Wallace, M. Jagielski, A. Herbert-Voss, K. Lee, A. Roberts, T. B. Brown, D. Song, U. Erlingsson, et al · 2021
Earlier work this paper cites.
Approximate data deletion from machine learning models
Z. Izzo, M. A. Smart, K. Chaudhuri, and J. Zou · 2021
Earlier work this paper cites.
Truthfulqa: Measuring how models mimic human falsehoods
S. Lin, J. Hilton, and O. Evans · 2021
Earlier work this paper cites.
Descent-to-delete: Gradient-based methods for machine unlearning
S. Neel, A. Roth, and S. Sharifi-Malvajerdi · 2021
Earlier work this paper cites.
Machine unlearning of features and labels
A. Warnecke, L. Pirch, C. Wressnegger, and K. Rieck · 2021
Cited alongside, same era.
Constitutional ai: Harmlessness from ai feedback
Y. Bai, S. Kadavath, S. Kundu, A. Askell, J. Kernion, A. Jones, A. Chen, A. Goldie, A. Mirhoseini, C. McKinnon, et al · 2022
Cited alongside, same era.
Quantifying memorization across neural language models
N. Carlini, D. Ippolito, M. Jagielski, K. Lee, F. Tramer, and C. Zhang · 2022
Cited alongside, same era.
Backdoor defense with machine unlearning
Y. Liu, M. Fan, C. Chen, X. Liu, Z. Ma, L. Wang, and J. Ma · 2022
Cited alongside, same era.
Quark: Controllable text generation with reinforced unlearning
X. Lu, S. Welleck, J. Hessel, L. Jiang, L. Qin, P. West, P. Ammanabrolu, and Y. Choi · 2022
Cited alongside, same era.
Model sparsification can simplify machine unlearning
J. Jia, J. Liu, P. Ram, Y. Yao, G. Liu, Y. Liu, P. Sharma, and S. Liu · 2023
Closest in time.
Do language models plagiarize?
J. Lee, T. Le, J. Chen, and D. Lee · 2023
Closest in time.
Halueval: A large-scale hallucination evaluation benchmark for large language models
J. Li, X. Cheng, W. X. Zhao, J.-Y. Nie, and J.-R. Wen · 2023
Closest in time.
Trustworthy llms: a survey and guideline for evaluating large language models’ alignment
Y. Liu, Y. Yao, J.-F. Ton, X. Zhang, R. G. H. Cheng, Y. Klochkov, M. F. Taufiq, and H. Li · 2023
Closest in time.
In-context unlearning: Language models as few shot unlearners
M. Pawelczyk, S. Neel, and H. Lakkaraju · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Probabilistic Machine Learning: An introduction
K. P. Murphy · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al · 2022
Cited alongside, same era.
Unrolling sgd: Understanding factors influencing machine unlearning
A. Thudi, G. Deza, V. Chandrasekaran, and N. Papernot · 2022
Cited alongside, same era.
How large language models are transforming machine-paraphrased plagiarism
J. P. Wahle, T. Ruas, F. Kirstein, and B. Gipp · 2022
Cited alongside, same era.
Opt: Open pre-trained transformer language models
S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin, et al · 2022
Cited alongside, same era.
Zero-shot machine unlearning
V. S. Chundawat, A. K. Tarun, M. Mandal, and M. Kankanhalli · 2023
Cited alongside, same era.
Github copilot lawsuit
G. Copilot · 2023
Cited alongside, same era.
Fast yet effective machine unlearning
A. K. Tarun, V. S. Chundawat, M. Mandal, and M. Kankanhalli · 2023
Closest in time.
Tiktok community guidelines
TikTok · 2023
Closest in time.
Llama 2: Open foundation and fine-tuned chat models
H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale, et al · 2023
Closest in time.
Twitter rules and policies
Twitter · 2023
Closest in time.
Machine unlearning: A survey
H. Xu, T. Zhu, L. Zhang, W. Zhou, and P. S. Yu · 2023
Closest in time.
Rrhf: Rank responses to align language models with human feedback without tears
Z. Yuan, H. Yuan, C. Tan, W. Wang, S. Huang, and F. Huang · 2023
Closest in time.
Secrets of rlhf in large language models part i: Ppo
R. Zheng, S. Dou, S. Gao, W. Shen, B. Wang, Y. Liu, S. Jin, Q. Liu, L. Xiong, L. Chen, et al · 2023
Closest in time.
T. Y. Zhuo, Z. Li, Y. Huang, Y.-F. Li, W. Wang, G. Haffari, and F. Shiri · 2023
Closest in time.
The times sues openai and microsoft over a.i. use of copyrighted work
M. Grynbaum and R. Mac · 2024
Closest in time.
Sarah silverman sues openai and meta over copyright infringement
Z. Small · 2024
Closest in time.