Fetching the paper…
Reading the bibliography…
The objective of digital forgetting is, given a model with undesirable knowledge or behavior, obtain a new model where the detected issues are no longer present.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. (2019) · 1905
Earlier work this paper cites.
Roberta: A robustly optimized BERT pretraining approach
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., and Stoyanov, V. (2019) · 1907
Earlier work this paper cites.
Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter
Sanh, V., Debut, L., Chaumond, J., and Wolf, T. (2019) · 1910
Earlier work this paper cites.
Samsum corpus: A human-annotated dialogue dataset for abstractive summarization
Gliwa, B., Mochol, I., Biesek, M., and Wawer, A. (2019) · 1911
Earlier work this paper cites.
Universal Declaration of Human Rights
United Nations (1948) · 1948
Earlier work this paper cites.
CrowS-pairs: A challenge dataset for measuring social biases in masked language models
Nangia, N., Vania, C., Bhalerao, R., and Bowman, S. R. (2020) · 1967
Earlier work this paper cites.
Harry Potter and the Sorcerer’s Stone
Rowling, J. K. (2000) · 2000
Earlier work this paper cites.
Introducing the enron corpus
Klimt, B. and Yang, Y. (2004) · 2004
Earlier work this paper cites.
Null it out: Guarding protected attributes by iterative nullspace projection
Ravfogel, S., Elazar, Y., Gonen, H., Twiton, M., and Goldberg, Y. (2020) · 2004
Earlier work this paper cites.
Learning word vectors for sentiment analysis
Maas, A. L., Daly, R. E., Pham, P. T., Huang, D., Ng, A. Y., and Potts, C. (2011) · 2011
Earlier work this paper cites.
SemEval-2012 task 7: Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Gordon, A., Kozareva, Z., and Roemmele, M. (2012) · 2012
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Mikolov, T., Chen, K., Corrado, G., and Dean, J. (2013) · 2013
Earlier work this paper cites.
Recursive deep models for semantic compositionality over a sentiment treebank
Socher, R., Perelygin, A., Wu, J., Chuang, J., Manning, C. D., Ng, A., and Potts, C. (2013) · 2013
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Bahdanau, D., Cho, K., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
Report on the 11th IWSLT evaluation campaign
Cettolo, M., Niehues, J., Stüker, S., Bentivogli, L., and Federico, M. (2014) · 2014
Earlier work this paper cites.
Learning phrase representations using rnn encoder-decoder for statistical machine translation
Cho, K., Van Merriënboer, B., Gulcehre, C., Bahdanau, D., Bougares, F., Schwenk, H., and Bengio, Y. (2014) · 2014
Earlier work this paper cites.
A gold standard dependency corpus for English
Silveira, N., Dozat, T., de Marneffe, M.-C., Bowman, S., Connor, M., Bauer, J., and Manning, C. (2014) · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Sutskever, I., Vinyals, O., and Le, Q. V. (2014) · 2014
Earlier work this paper cites.
Character-level convolutional networks for text classification
Zhang, X., Zhao, J. J., and LeCun, Y. (2015) · 2015
Earlier work this paper cites.
Deep learning with differential privacy
Abadi, M., Chu, A., Goodfellow, I., McMahan, H. B., Mironov, I., Talwar, K., and Zhang, L. (2016) · 2016
Earlier work this paper cites.
Regulation (EU) 2016/679 of the European Parliament and of the Council of 27 April 2016 on the protection of natural persons with regard to the processing of personal data and on the free movement of such data, and repealing Directive 95/46/EC (General Data Protection Regulation)
European Parliament and Council of the European Union (2016) · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Merity, S., Xiong, C., Bradbury, J., and Socher, R. (2016) · 2016
Earlier work this paper cites.
The LAMBADA dataset: Word prediction requiring a broad discourse context
Paperno, D., Kruszewski, G., Lazaridou, A., Pham, N. Q., Bernardi, R., Pezzelle, S., Baroni, M., Boleda, G., and Fernández, R. (2016) · 2016
Earlier work this paper cites.
C-sanitized: A privacy model for document redaction and sanitization
Sánchez, D. and Batet, M. (2016) · 2016
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D. (2017) · 2017
Earlier work this paper cites.
Zero-shot relation extraction via reading comprehension
Levy, O., Seo, M., Choi, E., and Zettlemoyer, L. (2017) · 2017
Earlier work this paper cites.
Membership inference attacks against machine learning models
Shokri, R., Stronati, M., Song, C., and Shmatikov, V. (2017) · 2017
Earlier work this paper cites.
Axiomatic attribution for deep networks
Sundararajan, M., Taly, A., and Yan, Q. (2017) · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I. (2017) · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the AI2 reasoning challenge
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O. (2018) · 2018
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. (2018) · 2018
Earlier work this paper cites.
Hierarchical neural story generation
Fan, A., Lewis, M., and Dauphin, Y. N. (2018) · 2018
Earlier work this paper cites.
Scalable private learning with pate
Papernot, N., Song, S., Mironov, I., Raghunathan, A., Talwar, K., and Erlingsson, Ú. (2018) · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., Sutskever, I., et al. (2018) · 2018
Earlier work this paper cites.
Gender bias in coreference resolution
Rudinger, R., Naradowsky, J., Leonard, B., and Van Durme, B. (2018) · 2018
Earlier work this paper cites.
Salem, A., Zhang, Y., Humbert, M., Berrang, P., Fritz, M., and Backes, M. (2018) · 2018
Earlier work this paper cites.
Mind the GAP: A balanced corpus of gendered ambiguous pronouns
Webster, K., Recasens, M., Axelrod, V., and Baldridge, J. (2018) · 2018
Earlier work this paper cites.
Privacy risk in machine learning: Analyzing the connection to overfitting
Yeom, S., Giacomelli, I., Fredrikson, M., and Jha, S. (2018) · 2018
Earlier work this paper cites.
Personalizing dialogue agents: I have a dog, do you have pets too?
Zhang, S., Dinan, E., Urbanek, J., Szlam, A., Kiela, D., and Weston, J. (2018) · 2018
Earlier work this paper cites.
Gender bias in coreference resolution: Evaluation and debiasing methods
Zhao, J., Wang, T., Yatskar, M., Ordonez, V., and Chang, K.-W. (2018) · 2018
Earlier work this paper cites.
MathQA: Towards interpretable math word problem solving with operation-based formalisms
Amini, A., Gabriel, S., Lin, S., Koncel-Kedziorski, R., Choi, Y., and Hajishirzi, H. (2019) · 2019
Earlier work this paper cites.
Nuanced metrics for measuring unintended bias with real data for text classification
Borkan, D., Dixon, L., Sorensen, J., Thain, N., and Vasserman, L. (2019) · 2019
Earlier work this paper cites.
Bias in bios: A case study of semantic representation bias in a high-stakes setting
De-Arteaga, M., Romanov, A., Wallach, H. M., Chayes, J. T., Borgs, C., Chouldechova, A., Geyik, S. C., Kenthapadi, K., and Kalai, A. T. (2019) · 2019
Earlier work this paper cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M., Lee, K., and Toutanova, K. (2019) · 2019
Cited alongside, same era.
Wizard of wikipedia: Knowledge-powered conversational agents
Dinan, E., Roller, S., Shuster, K., Fan, A., Auli, M., and Weston, J. (2019) · 2019
Cited alongside, same era.
Ethics guidelines for trustworthy AI
European Commission (2019) · 2019
Cited alongside, same era.
Automatic anonymization of textual documents: detecting sensitive information via word embeddings
Hassan, F., Sánchez, D., Soria-Comas, J., and Domingo-Ferrer, J. (2019) · 2019
Cited alongside, same era.
Parameter-efficient transfer learning for nlp
Houlsby, N., Giurgiu, A., Jastrzebski, S., Morrone, B., De Laroussilhe, Q., Gesmundo, A., Attariyan, M., and Gelly, S. (2019) · 2019
Cited alongside, same era.
Pubmedqa: A dataset for biomedical research question answering
Memorization without overfitting: Analyzing the training dynamics of large language models
Tirumala, K., Markosyan, A., Zettlemoyer, L., and Aghajanyan, A. (2022) · 2022
Later among the works it cites.
Natural language processing with transformers
Tunstall, L., Von Werra, L., and Wolf, T. (2022) · 2022
Later among the works it cites.
Opt: Open pre-trained transformer language models
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., Mihaylov, T., Ott, M., Shleifer, S., Shuster, K., Simig, D., Koura, P. S., Sridhar, A., Wang, T., and Zettlemoyer, L. (2022) · 2022
Later among the works it cites.
Leace: Perfect linear concept erasure in closed form
Belrose, N., Schneider-Joseph, D., Ravfogel, S., Cotterell, R., Raff, E., and Biderman, S. (2023) · 2023
Later among the works it cites.
Pythia: A suite for analyzing large language models across training and scaling
Biderman, S., Schoelkopf, H., Anthony, Q. G., Bradley, H., O’Brien, K., Hallahan, E., Khan, M. A., Purohit, S., Prashanth, U. S., Raff, E., et al. (2023) · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jin, Q., Dhingra, B., Liu, Z., Cohen, W. W., and Lu, X. (2019) · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al. (2019) · 2019
Cited alongside, same era.
Towards empathetic open-domain conversation models: A new benchmark and dataset
Rashkin, H., Smith, E. M., Li, M., and Boureau, Y.-L. (2019) · 2019
Cited alongside, same era.
Evaluating gender bias in machine translation
Stanovsky, G., Smith, N. A., and Zettlemoyer, L. (2019) · 2019
Cited alongside, same era.
PIQA: reasoning about physical commonsense in natural language
Bisk, Y., Zellers, R., Bras, R. L., Gao, J., and Choi, Y. (2020) · 2020
Cited alongside, same era.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D. (2020) · 2020
Cited alongside, same era.
ELECTRA: pre-training text encoders as discriminators rather than generators
Clark, K., Luong, M., Le, Q. V., and Manning, C. D. (2020) · 2020
Cited alongside, same era.
Later among the works it cites.
A critical review on the use (and misuse) of differential privacy in machine learning
Blanco-Justicia, A., Sánchez, D., Domingo-Ferrer, J., and Muralidhar, K. (2023) · 2023
Later among the works it cites.
What can we learn from data leakage and unlearning for law?
Borkar, J. (2023) · 2023
Later among the works it cites.
Unlearn what you want to forget: Efficient unlearning for llms
Chen, J. and Yang, D. (2023) · 2023
Later among the works it cites.
Redpajama: an open dataset for training large language models
Computer, T. (2023) · 2023
Later among the works it cites.
Elastic weight removal for faithful and abstractive dialogue generation
Daheim, N., Dziri, N., Sachan, M., Gurevych, I., and Ponti, E. M. (2023) · 2023
Later among the works it cites.
Who’s harry potter? approximate unlearning in llms
Eldan, R. and Russinovich, M. (2023) · 2023
Later among the works it cites.
Cook: Empowering general-purpose language models with modular and collaborative knowledge
Feng, S., Shi, W., Bai, Y., Balachandran, V., He, T., and Tsvetkov, Y. (2023) · 2023
Later among the works it cites.
Beavertails: Towards improved safety alignment of LLM via a human-preference dataset
Ji, J., Liu, M., Dai, J., Pan, X., Zhang, C., Bian, C., Zhang, B., Sun, R., Wang, Y., and Yang, Y. (2023) · 2023
Later among the works it cites.
Fairsisa: Ensemble post-processing to improve fairness of unlearning in llms
Kadhe, S., Halimi, A., Rawat, A., and Baracaldo, N. (2023) · 2023
Later among the works it cites.
Halueval: A large-scale hallucination evaluation benchmark for large language models
Li, J., Cheng, X., Zhao, X., Nie, J., and Wen, J. (2023a) · 2023
Later among the works it cites.
Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing
Liu, P., Yuan, W., Fu, J., Jiang, Z., Hayashi, H., and Neubig, G. (2023) · 2023
Later among the works it cites.
Forgetting private textual sequences in language models via leave-one-out ensemble
Liu, Z. and Kalinli, O. (2023) · 2023
Later among the works it cites.
Scalable extraction of training data from (production) language models
Nasr, M., Carlini, N., Hayase, J., Jagielski, M., Cooper, A. F., Ippolito, D., Choquette-Choo, C. A., Wallace, E., Tramèr, F., and Lee, K. (2023) · 2023
Later among the works it cites.
Ni, S., Chen, D., Li, C., Hu, X., Xu, R., and Yang, M. (2023) · 2023
Later among the works it cites.
Can sensitive information be deleted from llms? objectives for defending against extraction attacks
Patil, V., Hase, P., and Bansal, M. (2023) · 2023
Later among the works it cites.
In-context unlearning: Language models as few shot unlearners
Pawelczyk, M., Neel, S., and Lakkaraju, H. (2023) · 2023
Later among the works it cites.
Dissecting large language models
Pochinkov, N. and Schoots, N. (2023) · 2023
Later among the works it cites.
Learn to unlearn: A survey on machine unlearning
Qu, Y., Yuan, X., Ding, M., Ni, W., Rakotoarivelo, T., and Smith, D. (2023) · 2023
Later among the works it cites.
Exploring the landscape of machine unlearning: A survey and taxonomy
Shaik, T., Tao, X., Xie, H., Li, L., Zhu, X., and Li, Q. (2023) · 2023
Later among the works it cites.
Detecting pretraining data from large language models
Shi, W., Ajith, A., Xia, M., Huang, Y., Liu, D., Blevins, T., Chen, D., and Zettlemoyer, L. (2023) · 2023
Later among the works it cites.
Identifying and mitigating privacy risks stemming from language models: A survey
Smith, V., Shamsabadi, A. S., Ashurst, C., and Weller, A. (2023) · 2023
Later among the works it cites.
Kga: A general machine unlearning framework based on knowledge gap alignment
Wang, L., Chen, T., Yuan, W., Zeng, X., Wong, K.-F., and Yin, H. (2023) · 2023
Later among the works it cites.
Chain-of-thought prompting elicits reasoning in large language models
Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., and Zhou, D. (2023) · 2023
Later among the works it cites.
Depn: Detecting and editing privacy neurons in pretrained language models
Wu, X., Li, J., Xu, M., Dong, W., Wu, S., Bian, C., and Xiong, D. (2023) · 2023
Later among the works it cites.
Machine unlearning: A survey
Xu, H., Zhu, T., Zhang, L., Zhou, W., and Yu, P. S. (2023) · 2023
Later among the works it cites.
Large language model unlearning
Yao, Y., Xu, X., and Liu, Y. (2023) · 2023
Later among the works it cites.
Gradient ascent post-training enhances language model generalization
Yoon, D., Jang, J., Kim, S., and Seo, M. (2023) · 2023
Later among the works it cites.
Unlearning bias in language models by partitioning gradients
Yu, C., Jeoung, S., Kasi, A., Yu, P., and Ji, H. (2023a) · 2023
Later among the works it cites.
Right to be forgotten in the era of large language models: Implications, challenges, and solutions
Zhang, D., Finckenberg-Broman, P., Hoang, T., Pan, S., Xing, Z., Staples, M., and Xu, X. (2023) · 2023
Later among the works it cites.
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M. (2023) · 2023
Later among the works it cites.
Self-debiasing large language models: Zero-shot recognition and reduction of stereotypes
Gallegos, I. O., Rossi, R. A., Barrow, J., Tanjim, M. M., Yu, T., Deilamsalehy, H., Zhang, R., Kim, S., and Dernoncourt, F. (2024) · 2024
Closest in time.
Towards unbounded machine unlearning
Kurmanji, M., Triantafillou, P., Hayes, J., and Triantafillou, E. (2024) · 2024
Closest in time.
Tofu: A task of fictitious unlearning for llms
Maini, P., Feng, Z., Schwarzschild, A., Lipton, Z. C., and Kolter, J. Z. (2024) · 2024
Closest in time.
Aya model: An instruction finetuned open-access multilingual language model
Üstün, A., Aryabumi, V., Yong, Z.-X., Ko, W.-Y., D’souza, D., Onilude, G., Bhandari, N., Singh, S., Ooi, H.-L., Kayid, A., et al. (2024) · 2024
Closest in time.
Selective forgetting: Advancing machine unlearning techniques and evaluation in language models
Wang, L., Zeng, X., Guo, J., Wong, K.-F., and Gottlob, G. (2024) · 2024
Closest in time.
A comprehensive study of knowledge editing for large language models
Zhang, N., Yao, Y., Tian, B., Wang, P., Deng, S., Wang, M., Xi, Z., Mao, S., Zhang, J., Ni, Y., et al. (2024) · 2024
Closest in time.
Can you put it all together: Evaluating conversational agents’ ability to blend skills
Smith, E. M., Williamson, M., Shuster, K., Weston, J., and Boureau, Y.-L. (2020) · 2030
Closest in time.