Fetching the paper…
Reading the bibliography…
We present a novel benchmark and associated evaluation metrics for assessing the performance of text anonymization methods.
Scibert: A pretrained language model for scientific text
Beltagy, Iz, Kyle Lo, and Arman Cohan. 2019 · 1903
Earlier work this paper cites.
What privacy is for
Cohen, Julie E. 2012 · 1904
Earlier work this paper cites.
Roberta: A robustly optimized BERT pretraining approach
Liu, Yinhan, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019a · 1907
Earlier work this paper cites.
Privacy and Freedom
Westin, Alan F. 1967 · 1967
Earlier work this paper cites.
The european court of human rights and the protection of civil liberties: An overview
Gearty, Conor A. 1993 · 1993
Earlier work this paper cites.
Using information content to evaluate semantic similarity in a taxonomy
Resnik, Philip. 1995 · 1995
Earlier work this paper cites.
Replacing personally-identifying information in medical records, the scrub system
Sweeney, Latanya. 1996 · 1996
Earlier work this paper cites.
Protecting Privacy when Disclosing Information: k-Anonymity and its Enforcement through Generalization and Suppression
Samarati, Pierangela and Latanya Sweeney. 1998 · 1998
Earlier work this paper cites.
Annotating resources for information extraction
Boisen, Sean, Michael R. Crystal, Richard Schwartz, Rebecca Stone, and Ralph Weischedel. 2000 · 2000
Earlier work this paper cites.
Protecting respondents’ identities in microdata release
Samarati, Pierangela. 2001 · 2001
Earlier work this paper cites.
The Health Insurance Portability and Accountability Act
HIPAA. 2004 · 2004
Earlier work this paper cites.
Calibrating Noise to Sensitivity in Private Data Analysis
Dwork, Cynthia, Frank McSherry, Kobbi Nissim, and Adam Smith. 2006 · 2006
Earlier work this paper cites.
Revisiting the uniqueness of simple demographics in the US population
Golle, Philippe. 2006 · 2006
Earlier work this paper cites.
An introduction to NLP-based textual anonymisation
Medlock, Ben. 2006 · 2006
Earlier work this paper cites.
Privacy as a social good
Kasper, Debbie VS. 2007 · 2007
Earlier work this paper cites.
t-Closeness: Privacy Beyond k-Anonymity and l-Diversity
Li, Ninghui, Tiancheng Li, and Suresh Venkatasubramanian. 2007 · 2007
Earlier work this paper cites.
Web-Based Inference Detection
Staddon, Jessica, Philippe Golle, and Bryce Zimny. 2007 · 2007
Earlier work this paper cites.
Survey article: Inter-coder agreement for computational linguistics
Artstein, Ron and Massimo Poesio. 2008 · 2008
Earlier work this paper cites.
Efficient techniques for document sanitization
Chakaravarthy, Venkatesan T., Himanshu Gupta, Prasan Roy, and Mukesh K. Mohania. 2008 · 2008
Earlier work this paper cites.
Detecting privacy leaks using corpus-based association rules
Chow, Richard, Philippe Golle, and Jessica Staddon. 2008 · 2008
Earlier work this paper cites.
Automated de-identification of free-text medical records
Neamatullah, Ishna, Margaret M Douglass, H Lehman Li-wei, Andrew Reisner, Mauricio Villarroel, William J Long, Peter Szolovits, George B Moody, Roger G Mark, and Gari D Clifford. 2008 · 2008
Earlier work this paper cites.
The rules of redaction: Identify, protect, review (and repeat)
Bier, Eric A., Richard Chow, Philippe Golle, Tracy H. King, and J. Staddon. 2009 · 2009
Earlier work this paper cites.
Natural language processing with Python: analyzing text with the natural language toolkit
Bird, Steven, Ewan Klein, and Edward Loper. 2009 · 2009
Earlier work this paper cites.
Developing a standard for de-identifying electronic patient records written in Swedish: Precision, recall and f f -measure in a manual and computerized annotation trial
Velupillai, Sumithra, Hercules Dalianis, Martin Hassel, and Gunnar H. Nilsson. 2009 · 2009
Earlier work this paper cites.
The MITRE identification scrubber toolkit: design, training, and assessment
Aberdeen, John, Samuel Bayer, Reyyan Yeniterzi, Ben Wellner, Cheryl Clark, David Hanauer, Bradley Malin, and Lynette Hirschman. 2010 · 2010
Earlier work this paper cites.
Automatic de-identification of textual documents in the electronic health record: a review of recent research
Meystre, Stephane M, F Jeffrey Friedlin, Brett R South, Shuying Shen, and Matthew H Samore. 2010 · 2010
Earlier work this paper cites.
Significance of term relationships on anonymization
Anandan, Balamurugan and Chris Clifton. 2011 · 2011
Earlier work this paper cites.
A machine learning based system for semi-automatically redacting documents
Cumby, Chad M. and Rayid Ghani. 2011 · 2011
Earlier work this paper cites.
Text Classification for Data Loss Prevention
Hart, Michael, Pratyusa Manadhata, and Rob Johnson. 2011 · 2011
Earlier work this paper cites.
OntoNotes: A large training corpus for enhanced processing
Weischedel, Ralph, Eduard Hovy, Marcus. Mitchell, Palmer Martha S., Robert Belvin, Sameer S. Pradhan, Lance Ramshaw, and Nianwen Xue. 2011 · 2011
Earlier work this paper cites.
Pseudonymisation of personal names and other PHIs in an annotated clinical Swedish corpus
Alfalahi, Alyaa, Sara Brissman, and Hercules Dalianis. 2012 · 2012
Earlier work this paper cites.
t t -plausibility: Generalizing words to desensitize text
Anandan, Balamurugan, Chris Clifton, Wei Jiang, Mummoorthy Murugesan, Pedro Pastrana-Camacho, and Luo Si. 2012 · 2012
Earlier work this paper cites.
Evaluating current automatic de-identification methods with veteran’s health administration clinical documents
Ferrández, O., B. R. South, S. Shen, F. J. Friedlin, M. H. Samore, and S. M. Meystre. 2012 · 2012
Cited alongside, same era.
Statistical disclosure control
Hundepool, Anco, Josep Domingo-Ferrer, Luisa Franconi, Sarah Giessing, Eric Schulte Nordholt, Keith Spicer, and Peter-Paul De Wolf. 2012 · 2012
Cited alongside, same era.
Seven types of privacy
Finn, Rachel L, David Wright, and Michael Friedewald. 2013 · 2013
Cited alongside, same era.
Approaches of anonymisation of an SMS corpus
Patel, Namrata, Pierre Accorsi, Diana Inkpen, Cédric Lopez, and Mathieu Roche. 2013 · 2013
Cited alongside, same era.
Minimizing the disclosure risk of semantic correlations in document sanitization
Sánchez, David, Montserrat Batet, and Alexandre Viejo. 2013 · 2013
Cited alongside, same era.
The algorithmic foundations of differential privacy
A survey of automatic de-identification of longitudinal clinical narratives
Yogarajan, Vithya, Michael Mayo, and Bernhard Pfahringer. 2018 · 2018
Later among the works it cites.
Adversarial removal of demographic attributes revisited
Barrett, Maria, Yova Kementchedjhieva, Yanai Elazar, Desmond Elliott, and Anders Søgaard. 2019 · 2019
Later among the works it cites.
Towards private synthetic text generation
Bommasani, Rishi, Steven Wu, Zhiwei, and Alexandra K Schofield. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Devlin, Jacob, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Generalised differential privacy for text document processing
Fernandes, Natasha, Mark Dras, and Annabelle McIver. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dwork, Cynthia and Aaron Roth. 2014 · 2014
Cited alongside, same era.
Annotating longitudinal clinical narratives for de-identification: The 2014 i2b2/UTHealth corpus
Stubbs, Amber and Özlem Uzuner. 2015 · 2014
Cited alongside, same era.
Tm-score: A misuseability weight measure for textual content
Vartanian, Arik and Asaf Shabtai. 2014 · 2014
Cited alongside, same era.
Automatic anonymisation of a new portuguese-english parallel corpus in the legal-financial domains
Bick, Eckhard and Anabela Barreiro. 2015 · 2015
Cited alongside, same era.
Automatic detection of protected health information from clinic narratives
Yang, Hui and Jonathan M. Garibaldi. 2015 · 2015
Cited alongside, same era.
Demographic dialectal variation in social media: A case study of African-American English
Blodgett, Su Lin, Lisa Green, and Brendan O’Connor. 2016 · 2016
Cited alongside, same era.
Named entity recognition with bidirectional LSTM-CNNs
Chiu, Jason P.C. and Eric Nichols. 2016 · 2016
Cited alongside, same era.
Leveraging hierarchical representations for preserving privacy and utility in text
Feyisetan, Oluwaseyi, Tom Diethe, and Thomas Drake. 2019 · 2019
Later among the works it cites.
Decoupled weight decay regularization
Loshchilov, Ilya and Frank Hutter. 2019 · 2019
Later among the works it cites.
Automatic de-identification of medical texts in spanish: the meddocan track, corpus, guidelines, methods and evaluation of results
Marimon, Montserrat, Aitor Gonzalez-Agirre, Ander Intxaurrondo, Heidy Rodriguez, Jose Lopez Martin, Marta Villegas, and Martin Krallinger. 2019 · 2019
Later among the works it cites.
Deep reinforcement learning-based text anonymization against private-attribute inference
Mosallanezhad, Ahmadreza, Ghazaleh Beigi, and Huan Liu. 2019 · 2019
Later among the works it cites.
The value of protecting privacy
Santanen, Eric. 2019 · 2019
Later among the works it cites.
Privacy-aware text rewriting
Xu, Qiongkai, Lizhen Qu, Chenchen Xu, and Ran Cui. 2019 · 2019
Later among the works it cites.
Longformer: The long-document transformer
Beltagy, Iz, Matthew E. Peters, and Arman Cohan. 2020 · 2020
Later among the works it cites.
CodE Alltag 2.0 — a pseudonymized German-language email corpus
Eder, Elisabeth, Ulrike Krieg-Holz, and Udo Hahn. 2020 · 2020
Later among the works it cites.
spaCy: Industrial-strength Natural Language Processing in Python
Honnibal, Matthew, Ines Montani, Sofie Van Landeghem, and Adriane Boyd. 2020 · 2020
Later among the works it cites.
TextHide: Tackling data privacy in language understanding tasks
Huang, Yangsibo, Zhao Song, Danqi Chen, Kai Li, and Sanjeev Arora. 2020 · 2020
Later among the works it cites.
Deidentification of free-text medical records using pre-trained bidirectional transformers
Johnson, Alistair EW, Lucas Bulgarelli, and Tom J Pollard. 2020 · 2020
Later among the works it cites.
Biobert: a pre-trained biomedical language representation model for biomedical text mining
Lee, Jinhyuk, Wonjin Yoon, Sungdong Kim, Donghyeon Kim, Sunkyu Kim, Chan Ho So, and Jaewoo Kang. 2020 · 2020
Later among the works it cites.
Custom nlp approaches to data anonymization
Mendels, Omri. 2020 · 2020
Later among the works it cites.
Disruptive and avoidable: Gdpr challenges to secondary research uses of data
Peloquin, David, Michael DiMaio, Barbara Bierer, and Mark Barnes. 2020 · 2020
Later among the works it cites.
Anaphora and coreference resolution: A review
Sukthanker, Rhea, Soujanya Poria, Erik Cambria, and Ramkumar Thirunavukarasu. 2020 · 2020
Later among the works it cites.
When differential privacy meets NLP: The devil is in the detail
Habernal, Ivan. 2021 · 2021
Later among the works it cites.
Utility-preserving privacy protection of textual documents via word embeddings
Hassan, Fadi, David Sánchez, and Josep Domingo-Ferrer. 2021 · 2021
Later among the works it cites.
A privacy-preserving approach to extraction of personal information through automatic annotation and federated learning
Hathurusinghe, Rajitha, Isar Nejadgholi, and Miodrag Bolic. 2021 · 2021
Later among the works it cites.
De-identification of privacy-related entities in job postings
Jensen, Kristian Nørgaard, Mike Zhang, and Barbara Plank. 2021 · 2021
Later among the works it cites.
ADePT: Auto-encoder based differentially private text transformation
Krishna, Satyapriya, Rahul Gupta, and Christophe Dupuy. 2021 · 2021
Later among the works it cites.
Large language models can be strong differentially private learners
Li, Xuechen, Florian Tramèr, Percy Liang, and Tatsunori Hashimoto. 2021 · 2021
Later among the works it cites.
Anonymisation Models for Text Data: State of the Art, Challenges and Future Directions
Lison, Pierre, Ildikó Pilán, David Sánchez, Montserrat Batet, and Lilja Øvrelid. 2021 · 2021
Later among the works it cites.
No intruder, no validity: Evaluation criteria for privacy-preserving text anonymization
Mozes, Maximilian and Bennett Kleinberg. 2021 · 2021
Later among the works it cites.
Bootstrapping text anonymization models with distant supervision
Papadopoulou, Anthi, Pierre Lison, , Lilja Øvrelid, and Ildikó Pilán. 2022 · 2022
Closest in time.
The GDPR and unstructured data: is anonymization possible?
Weitzenboeck, Emily, Pierre Lison, Malgorzata Agnieszka Cyndecka, and Malcolm Langford. 2022 · 2022
Closest in time.
Leveraging synonymy and polysemy to improve semantic similarity assessments based on intrinsic information content
Batet, Montserrat and David Sánchez. 2020 · 2041
Closest in time.