Fetching the paper…
Reading the bibliography…
The potential of transformer-based LLMs risks being hindered by privacy concerns due to their reliance on extensive datasets, possibly including sensitive information.
“Calibrating noise to sensitivity in private data analysis”
Cynthia Dwork, Frank McSherry, Kobbi Nissim and Adam Smith · 2006
Earlier work this paper cites.
“The Composition Theorem for Differential Privacy”
Peter Kairouz, Sewoong Oh and Pramod Viswanath · 2015
Earlier work this paper cites.
“Regulation (EU) 2016/679 of the European Parliament and of the Council of 27 April 2016 on the protection of natural persons with regard to the processing of personal data and on the free movement of such data, and repealing Directive 95/46/EC (General Data Protection Regulation)”, 2016
European Parliament, European Council · 2016
Earlier work this paper cites.
“Pointer sentinel mixture models”
Stephen Merity, Caiming Xiong, James Bradbury and Richard Socher · 2016
Earlier work this paper cites.
“Membership Inference Attacks Against Machine Learning Models”
Reza Shokri, Marco Stronati, Congzheng Song and Vitaly Shmatikov · 2017
Earlier work this paper cites.
“Attention is all you need”
Ashish Vaswani et al · 2017
Earlier work this paper cites.
“The role of differential privacy in gdpr compliance”
Rachel Cummings and Deven Desai · 2018
Earlier work this paper cites.
“Bert: Pre-training of deep bidirectional transformers for language understanding”
Jacob Devlin, Ming-Wei Chang, Kenton Lee and Kristina Toutanova · 2018
Earlier work this paper cites.
“Improving language understanding by generative pre-training”
Alec Radford, Karthik Narasimhan, Tim Salimans and Ilya Sutskever · 2018
Earlier work this paper cites.
“California Consumer Privacy Act (CCPA)”, 2018
State of California · 2018
Cited alongside, same era.
“The secret sharer: Evaluating and testing unintended memorization in neural networks”
Nicholas Carlini et al · 2019
Cited alongside, same era.
“OpenWebText Corpus”, http://Skylion007.github.io/OpenWebTextCorpus , 2019
Aaron Gokaslan and Vanya Cohen · 2019
Cited alongside, same era.
“Language models are unsupervised multitask learners”
Alec Radford et al · 2019
Cited alongside, same era.
“BERT-based Lexical Substitution”
Wangchunshu Zhou et al · 2019
Cited alongside, same era.
“MetaPoison: Practical General-purpose Clean-label Data Poisoning”
W. Huang et al · 2020
Cited alongside, same era.
“Membership Inference Attacks From First Principles”
Nicholas Carlini et al · 2022
Later among the works it cites.
“Property inference from poisoning”
Saeed Mahloujifar, Esha Ghosh and Melissa Chase · 2022
Later among the works it cites.
“Truth serum: Poisoning machine learning models to reveal their secrets”
Florian Tramèr et al · 2022
Later among the works it cites.
“Extracting training data from diffusion models”
Nicolas Carlini et al · 2023
Later among the works it cites.
“Membership inference attacks against language models via neighbourhood comparison”
Justus Mattern et al · 2023
Later among the works it cites.
“Nationality Bias in Text Generation”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Extracting training data from large language models”
Nicholas Carlini et al · 2021
Cited alongside, same era.
“Concealed Data Poisoning Attacks on NLP Models”
Eric Wallace, Tony Zhao, Shi Feng and Sameer Singh · 2021
Cited alongside, same era.
“On the Importance of Difficulty Calibration in Membership Inference Attacks”
Lauren Watson, Chuan Guo, Graham Cormode and Alexandre Sablayrolles · 2021
Cited alongside, same era.
Pranav Narayanan et al · 2023
Later among the works it cites.
Jiashu Xu et al · 2023
Later among the works it cites.
“Virtual prompt injection for instruction-tuned large language models”
Jun Yan et al · 2023
Later among the works it cites.
“On the exploitability of instruction tuning”
Manli Shu et al · 2024
Closest in time.