Fetching the paper…
Reading the bibliography…
We investigate practical and scalable algorithms for training large language models (LLMs) with user-level differential privacy (DP) in order to provably safeguard all the examples contributed by each user.
Our data, ourselves: Privacy via distributed noise generation
Cynthia Dwork, Krishnaram Kenthapadi, Frank McSherry, Ilya Mironov, and Moni Naor · 2006
Earlier work this paper cites.
Beyond differential privacy: Composition theorems and relational logic for f-divergences between probabilistic programs
Gilles Barthe and Federico Olmedo · 2013
Earlier work this paper cites.
Deep Learning with Differential Privacy
Martin Abadi, Andy Chu, Ian Goodfellow, H. Brendan McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang · 2016
Earlier work this paper cites.
The complexity of differential privacy
Salil Vadhan · 2017
Earlier work this paper cites.
Differentially private federated learning: A client level perspective
Robin C Geyer, Tassilo Klein, and Moin Nabi · 2017
Earlier work this paper cites.
Communication-efficient learning of deep networks from decentralized data
Brendan McMahan, Eider Moore, Daniel Ramage, Seth Hampson, and Blaise Aguera y Arcas · 2017
Earlier work this paper cites.
Learning differentially private recurrent language models
H. Brendan McMahan, Daniel Ramage, Kunal Talwar, and Li Zhang · 2018
Earlier work this paper cites.
Privacy amplification by subsampling: tight analyses via couplings and divergences
Borja Balle, Gilles Barthe, and Marco Gaboardi · 2018
Earlier work this paper cites.
Gradient diversity: a key ingredient for scalable distributed learning
Dong Yin, Ashwin Pananjady, Max Lam, Dimitris Papailiopoulos, Kannan Ramchandran, and Peter Bartlett · 2018
Earlier work this paper cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern · 2018
Earlier work this paper cites.
cpSGD: Communication-efficient and differentially-private distributed SGD
Naman Agarwal, Ananda Theertha Suresh, Felix Xinnan X Yu, Sanjiv Kumar, and Brendan McMahan · 2018
Earlier work this paper cites.
Gmail Smart Compose: Real-Time Assisted Writing
Mia Xu Chen, Benjamin N Lee, Gagan Bansal, Yuan Cao, Shuyuan Zhang, Justin Lu, Jackie Tsay, Yinan Wang, Andrew M Dai, Zhifeng Chen, et al · 2019
Earlier work this paper cites.
The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks
Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, and Dawn Song · 2019
Earlier work this paper cites.
Auditing data provenance in text-generation models
Congzheng Song and Vitaly Shmatikov · 2019
Earlier work this paper cites.
TensorFlow Federated Stack Overflow dataset, 2019
The TensorFlow Federated Authors · 2019
Earlier work this paper cites.
Bounding User Contributions: A Bias-Variance Trade-off in Differential Privacy
Kareem Amin, Alex Kulesza, Andres Munoz, and Sergei Vassilvtiskii · 2019
Earlier work this paper cites.
Beyond Inferring Class Representatives: User-Level Privacy Leakage From Federated Learning
Zhibo Wang, Mengkai Song, Zhifei Zhang, Yang Song, Qian Wang, and Hairong Qi · 2019
Earlier work this paper cites.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Earlier work this paper cites.
Optimal private median estimation under minimal distributional assumptions
Christos Tzamos, Emmanouil-Vasileios Vlatakis-Gkaragkounis, and Ilias Zadik · 2020
Earlier work this paper cites.
Smoothly Bounding User Contributions in Differential Privacy
Alessandro Epasto, Mohammad Mahdian, Jieming Mao, Vahab Mirrokni, and Lijie Ren · 2020
Earlier work this paper cites.
Learning discrete distributions: user vs item-level privacy
Yuhan Liu, Ananda Theertha Suresh, Felix Xinnan X Yu, Sanjiv Kumar, and Michael Riley · 2020
Earlier work this paper cites.
Differentially private meta-learning
Jeffrey Li, Mikhail Khodak, Sebastian Caldas, and Ameet Talwalkar · 2020
Earlier work this paper cites.
Analyzing User-Level Privacy Attack Against Federated Learning
Mengkai Song, Zhibo Wang, Zhifei Zhang, Yang Song, Qian Wang, Ju Ren, and Hairong Qi · 2020
Earlier work this paper cites.
Auditing differentially private machine learning: How private is private SGD?
Matthew Jagielski, Jonathan Ullman, and Alina Oprea · 2020
Cited alongside, same era.
How many data points is a prompt worth?
Teven Le Scao and Alexander M Rush · 2021
Cited alongside, same era.
The power of scale for parameter-efficient 384 prompt tuning
Brian Lester, Rami Al-Rfou, and Noah Constant · 2021
Cited alongside, same era.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al · 2021
Cited alongside, same era.
Learning with User-Level privacy
Daniel Levy, Ziteng Sun, Kareem Amin, Satyen Kale, Alex Kulesza, Mehryar Mohri, and Ananda Theertha Suresh · 2021
Cited alongside, same era.
Quantifying memorization across neural language models
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang · 2023
Later among the works it cites.
User inference attacks on large language models
Nikhil Kandpal, Krishna Pillutla, Alina Oprea, Peter Kairouz, Christopher A Choquette-Choo, and Zheng Xu · 2023
Later among the works it cites.
User-level Private Stochastic Convex Optimization with Optimal Rates
Raef Bassily and Ziteng Sun · 2023
Later among the works it cites.
How to DP-fy ML: A Practical Guide to Machine Learning with Differential Privacy
Natalia Ponomareva, Hussein Hazimeh, Alex Kurakin, Zheng Xu, Carson Denison, H. Brendan McMahan, Sergei Vassilvitskii, Steve Chien, and Abhradeep Guha Thakurta · 2023
Later among the works it cites.
TAN Without a Burn: Scaling Laws of DP-SGD
Tom Sander, Pierre Stock, and Alexandre Sablayrolles · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
User-Level Privacy-Preserving Federated Learning: Analysis and Performance Optimization
Kang Wei, Jun Li, Ming Ding, Chuan Ma, Hang Su, Bo Zhang, and H Vincent Poor · 2021
Cited alongside, same era.
Tight differential privacy for discrete-valued mechanisms and for the subsampled gaussian mechanism using FFT
Antti Koskela, Joonas Jalko, Lukas Prediger, and Antti Honkela · 2021
Cited alongside, same era.
The Audio Auditor: User-Level Membership Inference in Internet of Things Voice Services
Yuantian Miao, Minhui Xue, Chao Chen, Lei Pan, Jun Zhang, Benjamin Zi Hao Zhao, Dali Kaafar, and Yang Xiang · 2021
Cited alongside, same era.
Unlocking high-accuracy differentially private image classification through scale
Soham De, Leonard Berrada, Jamie Hayes, Samuel L Smith, and Borja Balle · 2022
Cited alongside, same era.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Cited alongside, same era.
Deduplicating training data makes language models better
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini · 2022
Cited alongside, same era.
Improved differential privacy for SGD via optimal private linear operators on adaptive streams
Sergey Denisov, H Brendan McMahan, John Rush, Adam Smith, and Abhradeep Guha Thakurta · 2022
Cited alongside, same era.
Alexey Kurakin, Natalia Ponomareva, Umar Syed, Liam MacDermed, and Andreas Terzis · 2023
Later among the works it cites.
Towards federated foundation models: Scalable dataset pipelines for group-structured learning
Zachary Charles, Nicole Mitchell, Krishna Pillutla, Michael Reneer, and Zachary Garrett · 2023
Later among the works it cites.
Algorithms for bounding contribution for histogram estimation under user-level privacy
Yuhan Liu, Ananda Theertha Suresh, Wennan Zhu, Peter Kairouz, and Marco Gruteser · 2023
Later among the works it cites.
Continual Observation under User-level Differential Privacy
Wei Dong, Qiyao Luo, and Ke Yi · 2023
Later among the works it cites.
ULDP-FL: Federated learning with across silo user-level differential privacy
Fumiyuki Kato, Li Xiong, Shun Takagi, Yang Cao, and Masatoshi Yoshikawa · 2023
Later among the works it cites.
Federated Linear Contextual Bandits with User-level Differential Privacy
Ruiquan Huang, Huanyu Zhang, Luca Melis, Milan Shen, Meisam Hejazinia, and Jing Yang · 2023
Later among the works it cites.
Federated learning with differential privacy for end-to-end speech recognition
Martin Pelikan, Sheikh Shams Azam, Vitaly Feldman, Jan Silovsky, Kunal Talwar, Tatiana Likhomanenko, et al · 2023
Later among the works it cites.
Towards mobility reports with user-level privacy
Alexandra Kapp, Saskia Nuñez von Voigt, Helena Mihaljević, and Florian Tschorsch · 2023
Later among the works it cites.
FACE-AUDITOR: Data auditing in facial recognition systems
Min Chen, Zhikun Zhang, Tianhao Wang, Michael Backes, and Yang Zhang · 2023
Later among the works it cites.
Unleashing the Power of Randomization in Auditing Differentially Private ML
Krishna Pillutla, Galen Andrew, Peter Kairouz, H Brendan McMahan, Alina Oprea, and Sewoong Oh · 2023
Later among the works it cites.
Privacy Auditing with One (1) Training Run
Thomas Steinke, Milad Nasr, and Matthew Jagielski · 2023
Later among the works it cites.
User-level Differentially Private Stochastic Convex Optimization: Efficient Algorithms with Optimal Rates
Hilal Asi and Daogao Liu · 2024
Closest in time.
https://github.com/google/praxis
Praxis · 2024
Closest in time.
FAX: Scalable and differentiable federated primitives in jax
Keith Rush, Zachary Charles, and Zachary Garrett · 2024
Closest in time.
https://github.com/google/paxml
Paxml · 2024
Closest in time.
Continual Mean Estimation Under User-Level Privacy
Anand Jerry George, Lekshmi Ramesh, Aditya Vikram Singh, and Himanshu Tyagi · 2024
Closest in time.
Privacy Amplification by Sampling under User-level Differential Privacy
Juanru Fang and Ke Yi · 2024
Closest in time.
One-shot Empirical Privacy Estimation for Federated Learning
Galen Andrew, Peter Kairouz, Sewoong Oh, Alina Oprea, H Brendan McMahan, and Vinith Suriyakumar · 2024
Closest in time.