Fetching the paper…
Reading the bibliography…
Large language models (LLMs) have emerged as powerful tools for tackling complex tasks across diverse domains, but they also raise privacy concerns when fine-tuned on sensitive data due to potential memorization.
The Enron corpus: A new dataset for email classification research
Bryan Klimt and Yiming Yang · 2004
Earlier work this paper cites.
Deep learning with differential privacy
Martin Abadi, Andy Chu, Ian Goodfellow, H Brendan McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang · 2016
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas · 2017
Earlier work this paper cites.
The complexity of differential privacy
Salil P. Vadhan · 2017
Earlier work this paper cites.
Learning differentially private recurrent language models
H Brendan McMahan, Daniel Ramage, Kunal Talwar, and Li Zhang · 2018
Earlier work this paper cites.
Bounding user contributions: A bias-variance trade-off in differential privacy
Kareem Amin, Alex Kulesza, Andres Munoz, and Sergei Vassilvtiskii · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Beyond inferring class representatives: User-level privacy leakage from federated learning
Zhibo Wang, Mengkai Song, Zhifei Zhang, Yang Song, Qian Wang, and Hairong Qi · 2019
Earlier work this paper cites.
DP Accounting Library
Google’s DP Library · 2020
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Learning discrete distributions: user vs item-level privacy
Yuhan Liu, Ananda Theertha Suresh, Felix Xinnan X Yu, Sanjiv Kumar, and Michael Riley · 2020
Earlier work this paper cites.
Code for computing tight guarantees for differential privacy
Lukas Prediger and Antti Koskela · 2020
Earlier work this paper cites.
Differentially private SQL with bounded user contribution
Royce J Wilson, Celia Yuxin Zhang, William Lam, Damien Desfontaines, Daniel Simmons-Marengo, and Bryant Gipson · 2020
Earlier work this paper cites.
Differentially private medical texts generation using generative neural networks
Md Momin Al Aziz, Tanbir Ahmed, Tasnia Faequa, Xiaoqian Jiang, Yiyu Yao, and Noman Mohammed · 2021
Earlier work this paper cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al · 2021
Earlier work this paper cites.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al · 2021
Earlier work this paper cites.
User-level differentially private learning via correlated sampling
Badih Ghazi, Ravi Kumar, and Pasin Manurangsi · 2021
Earlier work this paper cites.
Numerical composition of differential privacy
Sivakanth Gopi, Yin Tat Lee, and Lukas Wutschitz · 2021
Earlier work this paper cites.
Practical and private (deep) learning without sampling or shuffling
Peter Kairouz, Brendan McMahan, Shuang Song, Om Thakkar, Abhradeep Thakurta, and Zheng Xu · 2021
Earlier work this paper cites.
Tight differential privacy for discrete-valued mechanisms and for the subsampled Gaussian mechanism using FFT
Antti Koskela, Joonas Jälkö, Lukas Prediger, and Antti Honkela · 2021
Cited alongside, same era.
Learning with user-level privacy
Daniel Levy, Ziteng Sun, Kareem Amin, Satyen Kale, Alex Kulesza, Mehryar Mohri, and Ananda Theertha Suresh · 2021
Cited alongside, same era.
A fast algorithm to optimally compose privacy guarantees of differentially private (DP) mechanisms to arbitrary accuracy
Microsoft · 2021
Cited alongside, same era.
Large scale private learning via low-rank reparametrization
Da Yu, Huishuai Zhang, Wei Chen, Jian Yin, and Tie-Yan Liu · 2021
Cited alongside, same era.
Large-scale differentially private BERT
Rohan Anil, Badih Ghazi, Vineet Gupta, Ravi Kumar, and Pasin Manurangsi · 2022
Cited alongside, same era.
Quantifying memorization across neural language models
Flocks of stochastic parrots: Differentially private prompt learning for large language models
Haonan Duan, Adam Dziedzic, Nicolas Papernot, and Franziska Boenisch · 2023
Later among the works it cites.
Legalbench: A collaboratively built benchmark for measuring legal reasoning in large language models
Neel Guha, Julian Nyarko, Daniel Ho, Christopher Ré, Adam Chilton, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel Rockmore, Diego Zambrano, et al · 2023
Later among the works it cites.
Foundation models and fair use
Peter Henderson, Xuechen Li, Dan Jurafsky, Tatsunori Hashimoto, Mark A Lemley, and Percy Liang · 2023
Later among the works it cites.
Privacy implications of retrieval-based language models
Yangsibo Huang, Samyak Gupta, Zexuan Zhong, Kai Li, and Danqi Chen · 2023
Later among the works it cites.
Benefits, limits, and risks of GPT-4 as an AI chatbot for medicine
Peter Lee, Sebastien Bubeck, and Joseph Petro · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang · 2022
Cited alongside, same era.
Connect the dots: Tighter discrete approximations of privacy loss distributions
Vadym Doroshenko, Badih Ghazi, Pritish Kamath, Ravi Kumar, and Pasin Manurangsi · 2022
Cited alongside, same era.
LoRA: Low-rank adaptation of large language models
Edward J Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen · 2022
Cited alongside, same era.
Booksum: A collection of datasets for long-form narrative summarization
Wojciech Kryściński, Nazneen Rajani, Divyansh Agarwal, Caiming Xiong, and Dragomir Radev · 2022
Cited alongside, same era.
Large language models can be strong differentially private learners
Xuechen Li, Florian Tramer, Percy Liang, and Tatsunori Hashimoto · 2022
Cited alongside, same era.
Biogpt: generative pre-trained transformer for biomedical text generation and mining
Renqian Luo, Liai Sun, Yingce Xia, Tao Qin, Sheng Zhang, Hoifung Poon, and Tie-Yan Liu · 2022
Cited alongside, same era.
Large language models encode clinical knowledge
Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al · 2022
Cited alongside, same era.
Starcoder: may the source be with you!
Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, et al · 2023
Later among the works it cites.
Foundation models for generalist medical artificial intelligence
Michael Moor, Oishi Banerjee, Zahra Shakeri Hossein Abad, Harlan M Krumholz, Jure Leskovec, Eric J Topol, and Pranav Rajpurkar · 2023
Later among the works it cites.
How to DP-fy ML: A practical guide to machine learning with differential privacy
Natalia Ponomareva, Hussein Hazimeh, Alex Kurakin, Zheng Xu, Carson Denison, H Brendan McMahan, Sergei Vassilvitskii, Steve Chien, and Abhradeep Guha Thakurta · 2023
Later among the works it cites.
Code Llama: Open foundation models for code
Baptiste Roziere, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, et al · 2023
Later among the works it cites.
Privacy-preserving in-context learning for large language models
Tong Wu, Ashwinee Panda, Jiachen T Wang, and Prateek Mittal · 2023
Later among the works it cites.
Learning to generate image embeddings with user-level differential privacy
Zheng Xu, Maxwell Collins, Yuxiao Wang, Liviu Panait, Sewoong Oh, Sean Augenstein, Ting Liu, Florian Schroff, and H Brendan McMahan · 2023
Later among the works it cites.
User-level differentially private stochastic convex optimization: Efficient algorithms with optimal rates
Hilal Asi and Daogao Liu · 2024
Closest in time.
Fine-tuning large language models with user-level differential privacy
Zachary Charles, Arun Ganesh, Ryan McKenna, H Brendan McMahan, Nicole Mitchell, Krishna Pillutla, and Keith Rush · 2024
Closest in time.
How private are DP-SGD implementations?
Lynn Chua, Badih Ghazi, Pritish Kamath, Ravi Kumar, Pasin Manurangsi, Amer Sinha, and Chiyuan Zhang · 2024
Closest in time.
Tight group-level DP guarantees for DP-SGD with sampling via mixture of Gaussians mechanisms
Arun Ganesh · 2024
Closest in time.
Deepseek-coder: When the large language model meets programming–the rise of code intelligence
Daya Guo, Qihao Zhu, Dejian Yang, Zhenda Xie, Kai Dong, Wentao Zhang, Guanting Chen, Xiao Bi, Y Wu, YK Li, et al · 2024
Closest in time.
Privacy-preserving in-context learning with differentially private few-shot generation
Xinyu Tang, Richard Shin, Huseyin A Inan, Andre Manoel, Fatemehsadat Mireshghallah, Zinan Lin, Sivakanth Gopi, Janardhan Kulkarni, and Robert Sim · 2024
Closest in time.
Synthetic text generation with differential privacy: A simple and practical recipe
Xiang Yue, Huseyin A Inan, Xuechen Li, Girish Kumar, Julia McAnallen, Hoda Shajari, Huan Sun, David Levitan, and Robert Sim · 2024
Closest in time.