Fetching the paper…
Reading the bibliography…
Studying data memorization in neural language models helps us understand the risks (e.g., to privacy or copyright) associated with models regurgitating training data and aids in the development of countermeasures.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Space/time trade-offs in hash coding with allowable errors
Burton H Bloom. 1970 · 1970
Earlier work this paper cites.
Weighted Bloom filter
Jehoshua Bruck, Jie Gao, and Anxiao Jiang. 2006 · 2006
Earlier work this paper cites.
Natural language processing with Python: analyzing text with the natural language toolkit
Steven Bird, Ewan Klein, and Edward Loper. 2009 · 2009
Earlier work this paper cites.
Training production language models without memorizing user data
Swaroop Ramaswamy, Om Thakkar, Rajiv Mathews, Galen Andrew, H Brendan McMahan, and Françoise Beaufays. 2020 · 2009
Earlier work this paper cites.
Comparison and evaluation of code clone detection techniques and tools: A qualitative approach
Chanchal K Roy, James R Cordy, and Rainer Koschke. 2009 · 2009
Earlier work this paper cites.
An evaluation framework for plagiarism detection
Martin Potthast, Benno Stein, Alberto Barrón-Cedeño, and Paolo Rosso. 2010 · 2010
Earlier work this paper cites.
Theory and practice of bloom filters for distributed systems
Sasu Tarkoma, Christian Esteve Rothenberg, and Eemil Lagerspetz. 2011 · 2011
Earlier work this paper cites.
Model inversion attacks that exploit confidence information and basic countermeasures
Matt Fredrikson, Somesh Jha, and Thomas Ristenpart. 2015 · 2015
Earlier work this paper cites.
Deep learning with differential privacy
Martin Abadi, Andy Chu, Ian Goodfellow, H Brendan McMahan, Ilya Mironov, Kunal Talwar, and Li Zhang. 2016 · 2016
Earlier work this paper cites.
A context-aware natural language generation dataset for dialogue systems
Ondrej Dušek and Filip Jurcıcek. 2016 · 2016
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Bloom filters, adaptivity, and the dictionary problem
Michael A Bender, Martin Farach-Colton, Mayank Goswami, Rob Johnson, Samuel McCauley, and Shikha Singh. 2018 · 2018
Earlier work this paper cites.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, and Dawn Song. 2019 · 2019
Earlier work this paper cites.
Semantic noise matters for neural natural language generation
Ondřej Dušek, David M Howcroft, and Verena Rieser. 2019 · 2019
Cited alongside, same era.
The Pile: An 800GB dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. 2020 · 2020
Cited alongside, same era.
Totto: A controlled table-to-text generation dataset
Ankur Parikh, Xuezhi Wang, Sebastian Gehrmann, Manaal Faruqui, Bhuwan Dhingra, Diyi Yang, and Dipanjan Das. 2020 · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, Peter J Liu, et al. 2020 · 2020
Cited alongside, same era.
Mlsum: The multilingual summarization corpus
Thomas Scialom, Paul-Alexis Dray, Sylvain Lamprier, Benjamin Piwowarski, and Jacopo Staiano. 2020 · 2020
Cited alongside, same era.
Differentially private fine-tuning of language models
Da Yu, Saurabh Naik, Arturs Backurs, Sivakanth Gopi, Huseyin A Inan, Gautam Kamath, Janardhan Kulkarni, Yin Tat Lee, Andre Manoel, Lukas Wutschitz, et al. 2021 · 2021
Later among the works it cites.
Counterfactual memorization in neural language models
Chiyuan Zhang, Daphne Ippolito, Katherine Lee, Matthew Jagielski, Florian Tramèr, and Nicholas Carlini. 2021 · 2021
Later among the works it cites.
Reconstructing training data with informed adversaries
Borja Balle, Giovanni Cherubin, and Jamie Hayes. 2022 · 2022
Closest in time.
What does it mean for a language model to preserve privacy?
Hannah Brown, Katherine Lee, Fatemehsadat Mireshghallah, Reza Shokri, and Florian Tramèr. 2022 · 2022
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The secret revealer: Generative model-inversion attacks against deep neural networks
Yuheng Zhang, Ruoxi Jia, Hengzhi Pei, Wenxiao Wang, Bo Li, and Dawn Song. 2020 · 2020
Cited alongside, same era.
Large-scale differentially private BERT
Rohan Anil, Badih Ghazi, Vineet Gupta, Ravi Kumar, and Pasin Manurangsi. 2021 · 2021
Cited alongside, same era.
Gpt-neo: Large scale autoregressive language modeling with mesh-tensorflow
Sid Black, Leo Gao, Phil Wang, Connor Leahy, and Stella Biderman. 2021 · 2021
Cited alongside, same era.
Extracting training data from large language models
Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, et al. 2021 · 2021
Cited alongside, same era.
Evaluating large language models trained on code
Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Ponde de Oliveira Pinto, Jared Kaplan, Harri Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, et al. 2021 · 2021
Cited alongside, same era.
The gem benchmark: Natural language generation, its evaluation and metrics
Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Anuoluwapo Aremu, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna Clinciu, Dipanjan Das, Kaustubh Dhole, et al. 2021 · 2021
Cited alongside, same era.
Deduplicating training data makes language models better
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. 2021 · 2021
Cited alongside, same era.
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, Katherine Lee, Florian Tramer, and Chiyuan Zhang. 2022 · 2022
Closest in time.
Palm: Scaling language modeling with pathways
Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, et al. 2022 · 2022
Closest in time.
About GitHub Copilot
GitHub. 2022 · 2022
Closest in time.
Reconstructing training data from trained neural networks
Niv Haim, Gal Vardi, Gilad Yehudai, Ohad Shamir, and Michal Irani. 2022 · 2022
Closest in time.
Deduplicating training data mitigates privacy risks in language models
Nikhil Kandpal, Eric Wallace, and Colin Raffel. 2022 · 2022
Closest in time.
Defending against reconstruction attacks with Rényi differential privacy
Pierre Stock, Igor Shilov, Ilya Mironov, and Alexandre Sablayrolles. 2022 · 2022
Closest in time.
Memorization without overfitting: Analyzing the training dynamics of large language models
Kushal Tirumala, Aram H Markosyan, Luke Zettlemoyer, and Armen Aghajanyan. 2022 · 2022
Closest in time.
Opt: Open pre-trained transformer language models
Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, et al. 2022 · 2022
Closest in time.
Provably confidential language modelling
Xuandong Zhao, Lei Li, and Yu-Xiang Wang. 2022 · 2022
Closest in time.
Pythia: A suite for analyzing large language models across training and scaling
Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al. 2023 · 2023
Closest in time.