Fetching the paper…
Reading the bibliography…
Dominant pre-trained language models (PLMs) have demonstrated the potential risk of memorizing and outputting the training data.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, et al. 2020 · 1901
Earlier work this paper cites.
Membership inference attacks from first principles
Nicholas Carlini, Steve Chien, Milad Nasr, et al. 2022 · 1914
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Mecab: Yet another part-of-speech and morphological analyzer
Taku Kudo. 2005 · 2005
Earlier work this paper cites.
A normalized levenshtein distance metric
Li Yujian and Liu Bo. 2007 · 2007
Earlier work this paper cites.
Newspaper paywalls—the hype and the reality
Merja Myllylahti. 2014 · 2014
Earlier work this paper cites.
Behind the newspaper paywall – lessons in charging for online content: a comparative analysis of why australian newspapers are stuck in the purgatorial space between digital and print
Andrea Carson. 2015 · 2015
Earlier work this paper cites.
Newspaper paywalls and corporate revenues: A comparative study
Merja Myllylahti. 2016 · 2016
Earlier work this paper cites.
ReCon: Revealing and controlling PII leaks in mobile network traffic
Jingjing Ren, Ashwin Rao, Martina Lindorfer, et al. 2016 · 2016
Earlier work this paper cites.
Introducing the paywall
Helle Sjøvaag. 2016 · 2016
Earlier work this paper cites.
Deep variational information bottleneck
Alexander A. Alemi, Ian Fischer, Joshua V. Dillon, et al. 2017 · 2017
Earlier work this paper cites.
Obfuscation-resilient privacy leak detection for mobile apps through differential analysis
Andrea Continella, Yanick Fratantonio, Martina Lindorfer, et al. 2017 · 2017
Earlier work this paper cites.
Implementation of a word segmentation dictionary called mecab-ipadic-neologd and study on how to use it effectively for information retrieval (in japanese)
Toshinori Sato, Taiichi Hashimoto, and Manabu Okumura. 2017 · 2017
Earlier work this paper cites.
Membership inference attacks against machine learning models
Reza Shokri, Marco Stronati, Congzheng Song, et al. 2017 · 2017
Earlier work this paper cites.
Subword regularization: Improving neural network translation models with multiple subword candidates
Taku Kudo. 2018 · 2018
Earlier work this paper cites.
Humans forget, machines remember: Artificial intelligence and the right to be forgotten
Tiffany Li, Eduard Fosch Villaronga, and Peter Kieseberg. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, et al. 2018 · 2018
Earlier work this paper cites.
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern. 2018 · 2018
Earlier work this paper cites.
The adverse effects of code duplication in machine learning models of code
Miltiadis Allamanis. 2019 · 2019
Earlier work this paper cites.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Nicholas Carlini, Chang Liu, Úlfar Erlingsson, et al. 2019 · 2019
Earlier work this paper cites.
Making AI forget you: data deletion in machine learning
Antonio A Ginart, Melody Y Guan, Gregory Valiant, et al. 2019 · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter. 2019 · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, et al. 2019 · 2019
Earlier work this paper cites.
The pile: An 800GB dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, et al. 2020 · 2020
Earlier work this paper cites.
Formalizing data deletion in the context of the right to be forgotten
Sanjam Garg, Shafi Goldwasser, and Prashant Nalini Vasudevan. 2020 · 2020
Cited alongside, same era.
Auditing differentially private machine learning: how private is private SGD?
Matthew Jagielski, Jonathan Ullman, and Alina Oprea. 2020 · 2020
Cited alongside, same era.
Newspapers’ content policy and the effect of paywalls on pageviews
Ho Kim, Reo Song, and Youngsoo Kim. 2020 · 2020
Cited alongside, same era.
KART: Parameterization of privacy leakage scenarios from pre-trained language models
Yuta Nakamura, Shouhei Hanaoka, Yukihiro Nomura, et al. 2020 · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Quantifying memorization across neural language models
Nicholas Carlini, Daphne Ippolito, Matthew Jagielski, et al. 2023 · 2023
Later among the works it cites.
On the creativity of large language models
Giorgio Franceschelli and Mirco Musolesi. 2023 · 2023
Later among the works it cites.
Exploring the limits of differentially private deep learning with group-wise clipping
Jiyan He, Xuechen Li, Da Yu, et al. 2023 · 2023
Later among the works it cites.
A VAE for transformers with nonparametric variational information bottleneck
James Henderson and Fabio James Fehr. 2023 · 2023
Later among the works it cites.
Preventing generation of verbatim memorization in language models gives a false sense of privacy
Daphne Ippolito, Florian Tramer, Milad Nasr, Chiyuan Zhang, Matthew Jagielski, Katherine Lee, Christopher Choquette Choo, and Nicholas Carlini. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
On the dangers of stochastic parrots: Can language models be too big?
Emily M Bender, Timnit Gebru, Angelina McMillan-Major, et al. 2021 · 2021
Cited alongside, same era.
On the opportunities and risks of foundation models
Rishi Bommasani, Drew A Hudson, Ehsan Adeli, et al. 2021 · 2021
Cited alongside, same era.
Extracting training data from large language models
Nicholas Carlini, Florian Tramèr, Eric Wallace, et al. 2021 · 2021
Cited alongside, same era.
Membership inference attack susceptibility of clinical language models
Abhyuday Jagannatha, Bhanu Pratap Singh Rawat, and Hong Yu. 2021 · 2021
Cited alongside, same era.
Does BERT pretrained on clinical notes reveal sensitive data?
Eric Lehman, Sarthak Jain, Karl Pichotta, Yoav Goldberg, and Byron Wallace. 2021 · 2021
Cited alongside, same era.
Adversary instantiation: Lower bounds for differentially private machine learning
Milad Nasr, Shuang Song, Abhradeep Thakurta, et al. 2021 · 2021
Cited alongside, same era.
Large scale private learning via low-rank reparametrization
Da Yu, Huishuai Zhang, Wei Chen, et al. 2021 · 2021
Cited alongside, same era.
Training data extraction from pre-trained language models: A survey
Shotaro Ishihara. 2023 · 2023
Later among the works it cites.
Stop uploading test data in plain text: Practical strategies for mitigating data contamination by evaluation benchmarks
Alon Jacovi, Avi Caciularu, Omer Goldman, and Yoav Goldberg. 2023 · 2023
Later among the works it cites.
Copyright violations and large language models
Antonia Karamolegkou, Jiaang Li, Li Zhou, and Anders Søgaard. 2023 · 2023
Later among the works it cites.
Do language models plagiarize?
Jooyoung Lee, Thai Le, Jinghui Chen, et al. 2023 · 2023
Later among the works it cites.
Membership inference attacks against language models via neighbourhood comparison
Justus Mattern, Fatemehsadat Mireshghallah, Zhijing Jin, Bernhard Schoelkopf, Mrinmaya Sachan, and Taylor Berg-Kirkpatrick. 2023 · 2023
Later among the works it cites.
How much do language models copy from their training data? evaluating linguistic novelty in text generation using RAVEN
R. Thomas McCoy, Paul Smolensky, Tal Linzen, Jianfeng Gao, and Asli Celikyilmaz. 2023 · 2023
Later among the works it cites.
Hanyin Shao, Jie Huang, Shen Zheng, et al. 2023 · 2023
Later among the works it cites.
Large language models encode clinical knowledge
Karan Singhal, Shekoofeh Azizi, Tao Tu, et al. 2023 · 2023
Later among the works it cites.
Bag of tricks for training data extraction from language models
Weichen Yu, Tianyu Pang, Qian Liu, et al. 2023 · 2023
Later among the works it cites.
Counterfactual memorization in neural language models
Chiyuan Zhang, Daphne Ippolito, Katherine Lee, et al. 2023 · 2023
Later among the works it cites.
Do membership inference attacks work on large language models?
Michael Duan, Anshuman Suri, Niloofar Mireshghallah, et al. 2024 · 2024
Closest in time.
DE-COP: Detecting copyrighted content in language models training data
André Vicente Duarte, Xuandong Zhao, Arlindo L. Oliveira, et al. 2024 · 2024
Closest in time.
Shotaro Ishihara. 2024 · 2024
Closest in time.
Sampling-based Pseudo-Likelihood for membership inference attacks
Masahiro Kaneko, Youmi Ma, Yuki Wata, and Naoaki Okazaki. 2024 · 2024
Closest in time.
Copyright traps for large language models
Matthieu Meeus, Igor Shilov, Manuel Faysse, et al. 2024 · 2024
Closest in time.
Detecting pretraining data from large language models
Weijia Shi, Anirudh Ajith, Mengzhou Xia, et al. 2024 · 2024
Closest in time.
Min-K%++: Improved baseline for detecting Pre-Training data from large language models
Jingyang Zhang, Jingwei Sun, Eric Yeats, Yang Ouyang, Martin Kuo, Jianyi Zhang, Hao Yang, and Hai Li. 2024 · 2024
Closest in time.
Are large pre-trained language models leaking your personal information?
Jie Huang, Hanyin Shao, and Kevin Chen-Chuan Chang. 2022 · 2047
Closest in time.