Fetching the paper…
Reading the bibliography…
The difficulty of anonymizing text data hinders the development and deployment of NLP in high-stakes domains that involve private data, such as healthcare and social services.
CTRL: a conditional transformer language model for controllable generation
Nitish Shirish Keskar, Bryan McCann, Lav R. Varshney, Caiming Xiong, and Richard Socher. 2019 · 1909
Earlier work this paper cites.
Physiobank, physiotoolkit, and physionet: components of a new research resource for complex physiologic signals
Ary L Goldberger, Luis AN Amaral, Leon Glass, Jeffrey M Hausdorff, Plamen Ch Ivanov, Roger G Mark, Joseph E Mietus, George B Moody, Chung-Kang Peng, and H Eugene Stanley. 2000 · 2000
Earlier work this paper cites.
Simple demographics often identify people uniquely
Latanya Sweeney. 2000 · 2000
Earlier work this paper cites.
Calibrating noise to sensitivity in private data analysis
Cynthia Dwork, Frank McSherry, Kobbi Nissim, and Adam Smith. 2006 · 2006
Earlier work this paper cites.
How does NLP benefit legal system: A summary of legal artificial intelligence
Haoxi Zhong, Chaojun Xiao, Cunchao Tu, Tianyang Zhang, Zhiyuan Liu, and Maosong Sun. 2020 · 2006
Earlier work this paper cites.
Robust de-anonymization of large sparse datasets
Arvind Narayanan and Vitaly Shmatikov. 2008 · 2008
Earlier work this paper cites.
Evaluating the state of the art in coreference resolution for electronic medical records
Ozlem Uzuner, Andreea Bodnari, Shuying Shen, Tyler Forbush, John Pestian, and Brett R South. 2012 · 2012
Earlier work this paper cites.
The algorithmic foundations of differential privacy
Cynthia Dwork, Aaron Roth, et al. 2014 · 2014
Earlier work this paper cites.
Deep learning with differential privacy
Martin Abadi, Andy Chu, Ian Goodfellow, H. Brendan McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang. 2016 · 2016
Earlier work this paper cites.
Automated anonymization of text documents
Nuno Mamede, Jorge Baptista, and Francisco Dias. 2016 · 2016
Earlier work this paper cites.
Invited paper: Local differential privacy on metric spaces: Optimizing the trade-off with utility
Mário Alvim, Konstantinos Chatzikokolakis, Catuscia Palamidessi, and Anna Pazii. 2018 · 2018
Earlier work this paper cites.
Hierarchical neural story generation
Angela Fan, Mike Lewis, and Yann Dauphin. 2018 · 2018
Earlier work this paper cites.
Generation of synthetic electronic medical record text
Jiaqi Guan, Runzhe Li, Sheng Yu, and Xuegong Zhang. 2018 · 2018
Earlier work this paper cites.
The secret sharer: Measuring unintended neural network memorization & extracting secrets
Nicholas Carlini, Chang Liu, Jernej Kos, Úlfar Erlingsson, and Dawn Song. 2019 · 2019
Earlier work this paper cites.
An empirical evaluation of deep learning for ICD-9 code assignment using MIMIC-III clinical notes
Jinmiao Huang, Cesar Osorio, and Luke Wicent Sy. 2019 · 2019
Earlier work this paper cites.
Evaluating semantic accuracy of data-to-text generation with natural language inference
Ondřej Dušek and Zdeněk Kasner. 2020 · 2020
Earlier work this paper cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020 · 2020
Earlier work this paper cites.
Towards privacy by design in learner corpora research: A case of on-the-fly pseudonymization of Swedish learner essays
Elena Volodina, Yousuf Ali Mohammed, Sandra Derbring, Arild Matsson, and Beata Megyesi. 2020 · 2020
Cited alongside, same era.
Generation and evaluation of privacy preserving synthetic health data
Andrew Yale, Saloni Dash, Ritik Dutta, Isabelle Guyon, Adrien Pavao, and Kristin P. Bennett. 2020 · 2020
Cited alongside, same era.
Differentially private medical texts generation using generative neural networks
Md Momin Al Aziz, Tanbir Ahmed, Tasnia Faequa, Xiaoqian Jiang, Yiyu Yao, and Noman Mohammed. 2021 · 2021
Cited alongside, same era.
The limits of differential privacy (and its misuse in data release and machine learning)
Josep Domingo-Ferrer, David Sánchez, and Alberto Blanco-Justicia. 2021 · 2021
Cited alongside, same era.
Trainable ranking models to evaluate the semantic accuracy of data-to-text neural generator
Nicolas Garneau and Luc Lamontagne. 2021 · 2021
Cited alongside, same era.
Adaptive differential privacy for language model training
Xinwei Wu, Li Gong, and Deyi Xiong. 2022 · 2022
Later among the works it cites.
Opacus: User-friendly differential privacy library in pytorch
Ashkan Yousefpour, Igor Shilov, Alexandre Sablayrolles, Davide Testuggine, Karthik Prasad, Mani Malek, John Nguyen, Sayan Ghosh, Akash Bharadwaj, Jessica Zhao, Graham Cormode, and Ilya Mironov. 2022 · 2022
Later among the works it cites.
Clinical text anonymization, its influence on downstream NLP tasks and the risk of re-identification
Iyadh Ben Cheikh Larbi, Aljoscha Burchardt, and Roland Roller. 2023 · 2023
Later among the works it cites.
A customized text sanitization mechanism with differential privacy
Sai Chen, Fengran Mo, Yanhao Wang, Cen Chen, Jian-Yun Nie, Chengyu Wang, and Jamie Cui. 2023 · 2023
Later among the works it cites.
MENLI: Robust evaluation metrics from natural language inference
Yanran Chen and Steffen Eger. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Numerical composition of differential privacy
Sivakanth Gopi, Yin Tat Lee, and Lukas Wutschitz. 2021 · 2021
Cited alongside, same era.
Coreference resolution without span representations
Yuval Kirstain, Ori Ram, and Omer Levy. 2021 · 2021
Cited alongside, same era.
Large language models can be strong differentially private learners
Xuechen Li, Florian Tramer, Percy Liang, and Tatsunori Hashimoto. 2021 · 2021
Cited alongside, same era.
Towards personal data anonymization for social messaging
Ondřej Sotolář, Jaromír Plhák, and David Šmahel. 2021 · 2021
Cited alongside, same era.
Differential privacy for text analytics via natural text sanitization
Xiang Yue, Minxin Du, Tianhao Wang, Yaliang Li, Huan Sun, and Sherman S. M. Chow. 2021 · 2021
Cited alongside, same era.
What does it mean for a language model to preserve privacy?
Hannah Brown, Katherine Lee, Fatemehsadat Mireshghallah, Reza Shokri, and Florian Tramèr. 2022 · 2022
Cited alongside, same era.
LoRA: Low-rank adaptation of large language models
Edward J Hu, yelong shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. 2022 · 2022
Cited alongside, same era.
Examining risks of racial biases in nlp tools for child protective services
Anjalie Field, Amanda Coston, Nupoor Gandhi, Alexandra Chouldechova, Emily Putnam-Hornstein, David Steier, and Yulia Tsvetkov. 2023 · 2023
Later among the works it cites.
Annotating mentions alone enables efficient domain adaptation for coreference resolution
Nupoor Gandhi, Anjalie Field, and Emma Strubell. 2023 · 2023
Later among the works it cites.
Harnessing large-language models to generate private synthetic text
Alexey Kurakin, Natalia Ponomareva, Umar Syed, Liam MacDermed, and Andreas Terzis. 2023 · 2023
Later among the works it cites.
Analyzing leakage of personally identifiable information in language models
N. Lukas, A. Salem, R. Sim, S. Tople, L. Wutschitz, and S. Zanella-Beguelin. 2023 · 2023
Later among the works it cites.
For generated text, is NLI-neutral text the best text?
Michail Mersinias and Kyle Mahowald. 2023 · 2023
Later among the works it cites.
Differentially private conditional text generation for synthetic data production
Pranav Putta, Ander Steele, and Joseph W Ferrara. 2023 · 2023
Later among the works it cites.
Privacy- and utility-preserving NLP with anonymized data: A case study of pseudonymization
Oleksandr Yermilov, Vipul Raheja, and Artem Chernodub. 2023 · 2023
Later among the works it cites.
Synthetic text generation with differential privacy: A simple and practical recipe
Xiang Yue, Huseyin Inan, Xuechen Li, Girish Kumar, Julia McAnallen, Hoda Shajari, Huan Sun, David Levitan, and Robert Sim. 2023 · 2023
Later among the works it cites.
Fine-tuning large language models with user-level differential privacy
Zachary Charles, Arun Ganesh, Ryan McKenna, H Brendan McMahan, Nicole Mitchell, Krishna Pillutla, and Keith Rush. 2024 · 2024
Closest in time.
Mind the privacy unit! User-level differential privacy for language model fine-tuning
Lynn Chua, Badih Ghazi, Yangsibo Huang, Pritish Kamath, Daogao Liu, Pasin Manurangsi, Amer Sinha, and Chiyuan Zhang. 2024 · 2024
Closest in time.
Sheared LLaMA: Accelerating language model pre-training via structured pruning
Mengzhou Xia, Tianyu Gao, Zhiyuan Zeng, and Danqi Chen. 2024 · 2024
Closest in time.