Fetching the paper…
Reading the bibliography…
Large Language Models (LLMs) have a privacy concern because they memorize training data (including personally identifiable information (PII) like emails and phone numbers) and leak it during inference.
The eu proposal for a general data protection regulation and the roots of the ‘right to be forgotten’
Mantelero, A · 2013
Earlier work this paper cites.
GloVe: Global vectors for word representation
Pennington, J., Socher, R., and Manning, C · 2014
Earlier work this paper cites.
Towards making systems forget with machine unlearning
Cao, Y. and Yang, J · 2015
Earlier work this paper cites.
https://eur-lex.europa.eu/eli/reg/2016/679/2016-05-04
Lex access to european union law · 2016
Earlier work this paper cites.
Pointer sentinel mixture models, 2016
Merity, S., Xiong, C., Bradbury, J., and Socher, R · 2016
Earlier work this paper cites.
https://www.priv.gc.ca/en/opc-news/news-and-announcements/2018/an_181010/ , October 2018
Privacy commissioner seeks federal court determination on key issue for canadians’ online reputation · 2018
Earlier work this paper cites.
Hierarchical neural story generation
Fan, A., Lewis, M., and Dauphin, Y · 2018
Earlier work this paper cites.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Carlini, N., Liu, C., Erlingsson, Ú., Kos, J., and Song, D · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K · 2019
Cited alongside, same era.
Machine unlearning, 2020
Bourtoule, L., Chandrasekaran, V., Choquette-Choo, C. A., Jia, H., Travers, A., Zhang, B., Lie, D., and Papernot, N · 2020
Cited alongside, same era.
Extracting training data from large language models
Carlini, N., Tramèr, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, Ú., Oprea, A., and Raffel, C · 2021
Cited alongside, same era.
Does BERT pretrained on clinical notes reveal sensitive data?
Lehman, E., Jain, S., Pichotta, K., Goldberg, Y., and Wallace, B · 2021
Cited alongside, same era.
Neuro-symbolic language modeling with automaton-augmented retrieval
Alon, U., Xu, F., He, J., Sengupta, S., Roth, D., and Neubig, G · 2022
Cited alongside, same era.
The privacy onion effect: Memorization is relative
Deduplicating training data mitigates privacy risks in language models, 2022
Kandpal, N., Wallace, E., and Raffel, C · 2022
Later among the works it cites.
An empirical analysis of memorization in fine-tuned autoregressive language models
Mireshghallah, F., Uniyal, A., Wang, T., Evans, D., and Berg-Kirkpatrick, T · 2022
Later among the works it cites.
Deduplicating training data makes language models better, 2022
Nystrom, A., Zhang, C., Callison-Burch, C., Ippolito, D., Eck, D., Lee, K., and Carlini, N · 2022
Later among the works it cites.
Emergent and predictable memorization in large language models, 2023
Biderman, S., Prashanth, U. S., Sutawika, L., Schoelkopf, H., Anthony, Q., Purohit, S., and Raf, E · 2023
Closest in time.
Quantifying memorization across neural language models, 2023
Carlini, N., Ippolito, D., Jagielski, M., Lee, K., Tramer, F., and Zhang, C · 2023
Closest in time.
Analyzing leakage of personally identifiable information in language models, 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Carlini, N., Jagielski, M., Zhang, C., Papernot, N., Terzis, A., and Tramer, F · 2022
Cited alongside, same era.
Are large pre-trained language models leaking your personal information?
Huang, J., Shao, H., and Chang, K. C.-C · 2022
Cited alongside, same era.
https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=201720180AB375
Bill text
Cited in the paper.
https://huggingface.co/datasets/jaydeepb/wiki2emailsdataset
Wikitext-2 emails dataset
Cited in the paper.
https://github.com/RaRe-Technologies/gensim-data
gensim-data
Cited in the paper.
https://huggingface.co/jaydeepb/gpt2-wiki-emails-no-pattern
Wikitext-2 emails gpt-2
Cited in the paper.
https://www.zlib.net/
Zlib compression library
Cited in the paper.
Lukas, N., Salem, A., Sim, R., Tople, S., Wutschitz, L., and Zanella-Béguelin, S · 2023
Closest in time.
Machine learning model attribution challenge
Merkhofer, E., Chaudhari, D., Anderson, H. S., Manville, K., Wong, L., and Gante, J · 2023
Closest in time.