Fetching the paper…
Reading the bibliography…
Generative language models (LMs) offer numerous advantages but may produce inappropriate or harmful outputs due to the harmful knowledge acquired during pre-training.
Mathqa: Towards interpretable math word problem solving with operation-based formalisms
A. Amini, S. Gabriel, S. Lin, R. Koncel-Kedziorski, Y. Choi, and H. Hajishirzi · 1905
Earlier work this paper cites.
Scanning electronic documents for personally identifiable information
T. Aura, T. A. Kuhn, and M. Roe · 2006
Earlier work this paper cites.
Differential Privacy: A Survey of Results , page 1–19
C. Dwork · 2008
Earlier work this paper cites.
SemEval-2012 task 7: Choice of plausible alternatives: An evaluation of commonsense causal reasoning
A. Gordon, Z. Kozareva, and M. Roemmele · 2012
Earlier work this paper cites.
Deep learning with differential privacy
M. Abadi, A. Chu, I. Goodfellow, H. B. McMahan, I. Mironov, K. Talwar, and L. Zhang · 2016
Earlier work this paper cites.
De-identification of patient notes with recurrent neural networks
F. Dernoncourt, J. Y. Lee, O. Uzuner, and P. Szolovits · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
D. Paperno, G. Kruszewski, A. Lazaridou, N. Q. Pham, R. Bernardi, S. Pezzelle, M. Baroni, G. Boleda, and R. Fernández · 2016
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
P. Clark, I. Cowhey, O. Etzioni, T. Khot, A. Sabharwal, C. Schoenick, and O. Tafjord · 2018
Earlier work this paper cites.
Algorithms that remember: model inversion attacks and data protection law
M. Veale, R. Binns, and L. Edwards · 2018
Earlier work this paper cites.
Unsupervised feature learning via non-parametric instance discrimination
Z. Wu, Y. Xiong, S. X. Yu, and D. Lin · 2018
Earlier work this paper cites.
PubMedQA: A dataset for biomedical research question answering
Q. Jin, B. Dhingra, Z. Liu, W. Cohen, and X. Lu · 2019
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale
K. Sakaguchi, R. L. Bras, C. Bhagavatula, and Y. Choi · 2019
Earlier work this paper cites.
HellaSwag: Can a machine really finish your sentence?
R. Zellers, A. Holtzman, Y. Bisk, A. Farhadi, and Y. Choi · 2019
Cited alongside, same era.
Piqa: Reasoning about physical commonsense in natural language
Y. Bisk, R. Zellers, R. Le bras, J. Gao, and Y. Choi · 2020
Cited alongside, same era.
Momentum contrast for unsupervised visual representation learning
K. He, H. Fan, Y. Wu, S. Xie, and R. Girshick · 2020
Cited alongside, same era.
nlp-fluency
J. Bao · 2021
Cited alongside, same era.
Gpt-neo: Large scale autoregressive language modeling with mesh-tensorflow
S. Black, L. Gao, P. Wang, C. Leahy, and S. Biderman · 2021
Cited alongside, same era.
SimCSE: Simple contrastive learning of sentence embeddings
T. Gao, X. Yao, and D. Chen · 2021
Cited alongside, same era.
Knowledge unlearning for mitigating privacy risks in language models
J. Jang, D. Yoon, S. Yang, S. Cha, M. Lee, L. Logeswaran, and M. Seo · 2023
Later among the works it cites.
Beavertails: Towards improved safety alignment of llm via a human-preference dataset
J. Ji, M. Liu, J. Dai, X. Pan, C. Zhang, C. Bian, C. Zhang, R. Sun, Y. Wang, and Y. Yang · 2023
Later among the works it cites.
Towards unbounded machine unlearning
M. Kurmanji, P. Triantafillou, J. Hayes, and E. Triantafillou · 2023
Later among the works it cites.
Multi-step jailbreaking privacy attacks on chatgpt
H. Li, D. Guo, W. Fan, M. Xu, and Y. Song · 2023
Later among the works it cites.
In-context unlearning: Language models as few shot unlearners
M. Pawelczyk, S. Neel, and H. Lakkaraju · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Efficient two-stage model retraining for machine unlearning
J. Kim and S. S. Woo · 2022
Cited alongside, same era.
Sgpt: Gpt sentence embeddings for semantic search
N. Muennighoff · 2022
Cited alongside, same era.
Opt: Open pre-trained transformer language models
S. Zhang, S. Roller, N. Goyal, M. Artetxe, M. Chen, S. Chen, C. Dewan, M. Diab, X. Li, X. V. Lin, et al · 2022
Cited alongside, same era.
Unlearn what you want to forget: Efficient unlearning for LLMs
J. Chen and D. Yang · 2023
Cited alongside, same era.
Who’s harry potter? approximate unlearning in llms
R. Eldan and M. Russinovich · 2023
Cited alongside, same era.
KGA: A general machine unlearning framework based on knowledge gap alignment
L. Wang, T. Chen, W. Yuan, X. Zeng, K.-F. Wong, and H. Yin · 2023
Later among the works it cites.
Large language model unlearning
Y. Yao, X. Xu, and Y. Liu · 2023
Later among the works it cites.
Removing rlhf protections in gpt-4 via fine-tuning
Q. Zhan, R. Fang, R. Bindu, A. Gupta, T. Hashimoto, and D. Kang · 2023
Later among the works it cites.
Audit to forget: A unified method to revoke patients’ private data in intelligent healthcare
J. Zhou, H. Li, X. Liao, B. Zhang, W. He, Z. Li, L. Zhou, and X. Gao · 2023
Later among the works it cites.
Debiasing machine unlearning with counterfactual examples
Z. Chen, J. Wang, J. Zhuang, A. G. Reddy, F. Silvestri, J. Huang, K. Nag, K. Kuang, X. Ning, and G. Tolomei · 2024
Closest in time.
Contrastive unlearning: A contrastive approach to machine unlearning
H. kyu Lee, Q. Zhang, C. Yang, J. Lou, and L. Xiong · 2024
Closest in time.