Fetching the paper…
Reading the bibliography…
Past work has shown that large language models are susceptible to privacy attacks, where adversaries generate sequences from a trained model and detect which sequences are memorized from the training set.
Calibrating noise to sensitivity in private data analysis
Dwork, C., McSherry, F., Nissim, K., and Smith, A · 2006
Earlier work this paper cites.
Model inversion attacks that exploit confidence information and basic countermeasures
Fredrikson, M., Jha, S., and Ristenpart, T · 2015
Earlier work this paper cites.
Semeval-2017 task 1: Semantic textual similarity multilingual and crosslingual focused evaluation
Cer, D., Diab, M., Agirre, E., Lopez-Gazpio, I., and Specia, L · 2017
Earlier work this paper cites.
Model inversion attacks for prediction systems: Without knowledge of non-sensitive attributes
Hidano, S., Murakami, T., Katsumata, S., Kiyomoto, S., and Hanaoka, G · 2017
Earlier work this paper cites.
Membership inference attacks against machine learning models
Shokri, R., Stronati, M., Song, C., and Shmatikov, V · 2017
Earlier work this paper cites.
Attention is all you need
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., and Polosukhin, I · 2017
Earlier work this paper cites.
Privacy-preserving neural representations of text
Coavoux, M., Narayan, S., and Cohen, S. B · 2018
Earlier work this paper cites.
Hierarchical neural story generation
Fan, A., Lewis, M., and Dauphin, Y · 2018
Earlier work this paper cites.
Towards robust and privacy-preserving text representations
Li, Y., Baldwin, T., and Cohn, T · 2018
Earlier work this paper cites.
Understanding membership inferences on well-generalized learning models
Long, Y., Bindschaedler, V., Wang, L., Bu, D., Wang, X., Tang, H., Gunter, C. A., and Chen, K · 2018
Earlier work this paper cites.
Do CIFAR-10 classifiers generalize to CIFAR-10?
Recht, B., Roelofs, R., Schmidt, L., and Shankar, V · 2018
Earlier work this paper cites.
Privacy risk in machine learning: Analyzing the connection to overfitting
Yeom, S., Giacomelli, I., Fredrikson, M., and Jha, S · 2018
Earlier work this paper cites.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Carlini, N., Liu, C., Erlingsson, Ú., Kos, J., and Song, D · 2019
Earlier work this paper cites.
OpenWebText corpus, 2019
Gokaslan, A., Cohen, V., Pavlick, E., and Tellex, S · 2019
Earlier work this paper cites.
Language models as knowledge bases?
Petroni, F., Rocktäschel, T., Riedel, S., Lewis, P., Bakhtin, A., Wu, Y., and Miller, A · 2019
Cited alongside, same era.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Cited alongside, same era.
Auditing data provenance in text-generation models
Song, C. and Shmatikov, V · 2019
Cited alongside, same era.
Neural network inversion in adversarial setting via background knowledge alignment
Yang, Z., Zhang, J., Chang, E.-C., and Liang, Z · 2019
Cited alongside, same era.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2020
Cited alongside, same era.
Does BERT pretrained on clinical notes reveal sensitive data?
Lehman, E., Jain, S., Pichotta, K., Goldberg, Y., and Wallace, B · 2021
Later among the works it cites.
McCoy, R. T., Smolensky, P., Linzen, T., Gao, J., and Celikyilmaz, A · 2021
Later among the works it cites.
Privacy regularization: Joint privacy-utility optimization in LanguageModels
Mireshghallah, F., Inan, H., Hasegawa, M., Rühle, V., Berg-Kirkpatrick, T., and Sim, R · 2021
Later among the works it cites.
On memorization in probabilistic deep generative models
Van den Burg, G. and Williams, C · 2021
Later among the works it cites.
On the importance of difficulty calibration in membership inference attacks
Watson, L., Guo, C., Cormode, G., and Sablayrolles, A · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
What neural networks memorize and why: Discovering the long tail via influence estimation
Feldman, V. and Zhang, C · 2020
Cited alongside, same era.
The Pile: An 800GB dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Cited alongside, same era.
How much knowledge can you pack into the parameters of a language model?
Roberts, A., Raffel, C., and Shazeer, N · 2020
Cited alongside, same era.
Information Leakage in Embedding Models , pp. 377–390
Song, C. and Raghunathan, A · 2020
Cited alongside, same era.
When is memorization of irrelevant training data necessary for high-accuracy learning?
Brown, G., Bun, M., Feldman, V., Smith, A., and Talwar, K · 2021
Cited alongside, same era.
Training data leakage analysis in language models
Inan, H. A., Ramadan, O., Wutschitz, L., Jones, D., Rühle, V., Withers, J., and Sim, R · 2021
Cited alongside, same era.
Reflective decoding: Beyond unidirectional generation with off-the-shelf language models
West, P., Lu, X., Holtzman, A., Bhagavatula, C., Hwang, J., and Choi, Y · 2021
Later among the works it cites.
Differentially private fine-tuning of language models
Yu, D., Naik, S., Backurs, A., Gopi, S., Inan, H. A., Kamath, G., Kulkarni, J., Lee, Y. T., Manoel, A., Wutschitz, L., et al · 2021
Later among the works it cites.
Counterfactual memorization in neural language models
Zhang, C., Ippolito, D., Lee, K., Jagielski, M., Tramèr, F., and Carlini, N · 2021
Later among the works it cites.
A first look at rote learning in GitHub Copilot suggestions, June 2021
Ziegler, A · 2021
Later among the works it cites.
Quantifying memorization across neural language models
Carlini, N., Ippolito, D., Jagielski, M., Lee, K., Tramer, F., and Zhang, C · 2022
Closest in time.
Scaling laws and interpretability of learning from repeated data
Hernandez, D., Brown, T., Conerly, T., DasSarma, N., Drain, D., El-Showk, S., Elhage, N., Hatfield-Dodds, Z., Henighan, T., Hume, T., Johnston, S., Mann, B., Olah, C., Olsson, C., Amodei, D., Joseph, N., Kaplan, J., and McCandlish, S · 2022
Closest in time.
Large language models can be strong differentially private learners
Li, X., Tramèr, F., Liang, P., and Hashimoto, T · 2022
Closest in time.
Provably confidential language modelling
Zhao, X., Li, L., and Wang, Y.-X · 2022
Closest in time.