Fetching the paper…
Reading the bibliography…
This paper studies extractable memorization: training data that an adversary can efficiently extract by querying a machine learning model without prior knowledge of the training dataset.
The population frequencies of species and the estimation of population parameters
Good, I. J · 1953
Earlier work this paper cites.
Label-only membership inference attacks
Choquette-Choo, C. A., Tramer, F., Carlini, N., and Papernot, N · 1974
Earlier work this paper cites.
Nonparametric estimation of the number of classes in a population
Chao, A · 1984
Earlier work this paper cites.
Smooth nonparametric estimation of the quantile function
Zelterman, D · 1990
Earlier work this paper cites.
Good-Turing frequency estimation without tears
Gale, W. A., and Sampson, G · 1995
Earlier work this paper cites.
Ecological methods
Southwood, T. R. E., and Henderson, P. A · 2009
Earlier work this paper cites.
An improved nonparametric lower bound of species richness via a modified good–turing frequency formula
Chiu, C.-H., Wang, Y.-T., Walther, B. A., and Chao, A · 2014
Earlier work this paper cites.
Model inversion attacks that exploit confidence information and basic countermeasures
Fredrikson, M., Jha, S., and Ristenpart, T · 2015
Earlier work this paper cites.
Deep reinforcement learning from human preferences
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D · 2017
Earlier work this paper cites.
Membership inference attacks against machine learning models
Shokri, R., Stronati, M., Song, C., and Shmatikov, V · 2017
Earlier work this paper cites.
Privacy risk in machine learning: Analyzing the connection to overfitting
Yeom, S., Giacomelli, I., Fredrikson, M., and Jha, S · 2018
Earlier work this paper cites.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Carlini, N., Liu, C., Erlingsson, Ú., Kos, J., and Song, D · 2019
Earlier work this paper cites.
Language Models are Unsupervised Multitask Learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., and Sutskever, I · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., et al · 2020
Earlier work this paper cites.
The Pile: An 800GB dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., et al · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 2020
Earlier work this paper cites.
GPT-Neo: Large scale autoregressive language modeling with Mesh-Tensorflow, 2021
Black, S., Gao, L., Wang, P., Leahy, C., and Biderman, S · 2021
Earlier work this paper cites.
Extracting training data from large language models
Carlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, U., et al · 2021
Earlier work this paper cites.
Vulnerability disclosure policy
Project Zero · 2021
Cited alongside, same era.
Multitask prompted training enables zero-shot task generalization
Sanh, V., Webson, A., Raffel, C., Bach, S. H., Sutawika, L., Alyafeai, Z., Chaffin, A., Stiegler, A., Scao, T. L., Raja, A., et al · 2021
Cited alongside, same era.
Github Copilot research recitation, 2021
Ziegler, A · 2021
Cited alongside, same era.
Training a helpful and harmless assistant with reinforcement learning from human feedback
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al · 2022
Cited alongside, same era.
Reconstructing training data with informed adversaries
Balle, B., Cherubin, G., and Hayes, J · 2022
Cited alongside, same era.
What does it mean for a language model to preserve privacy?
Are aligned neural networks adversarially aligned?
Carlini, N., Nasr, M., Choquette-Choo, C. A., Jagielski, M., Gao, I., Awadalla, A., Koh, P. W., Ippolito, D., Lee, K., Tramer, F., et al · 2023
Closest in time.
RedPajama: An open source recipe to reproduce LLaMA training dataset, 2023
Computer, T · 2023
Closest in time.
Releasing 3B and 7B RedPajama-INCITE family of models including base, instruction-tuned & chat models, 2023
Computer, T · 2023
Closest in time.
Training data extraction from pre-trained language models: A survey, 2023
Ishihara, S · 2023
Closest in time.
Mistral 7b, 2023
Jiang, A. Q., Sablayrolles, A., Mensch, A., Bamford, C., Chaplot, D. S., de las Casas, D., Bressand, F., Lengyel, G., Lample, G., Saulnier, L., Lavaud, L. R., Lachaux, M.-A., Stock, P., Scao, T. L., Lavril, T., Wang, T., Lacroix, T., and Sayed, W. E · 2023
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Brown, H., Lee, K., Mireshghallah, F., Shokri, R., and Tramèr, F · 2022
Cited alongside, same era.
Membership inference attacks from first principles
Carlini, N., Chien, S., Nasr, M., Song, S., Terzis, A., and Tramer, F · 2022
Cited alongside, same era.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Cited alongside, same era.
An empirical analysis of compute-optimal large language model training
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de Las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Cited alongside, same era.
Deduplicating training data mitigates privacy risks in language models
Kandpal, N., Wallace, E., and Raffel, C · 2022
Cited alongside, same era.
Deduplicating training data makes language models better
Lee, K., Ippolito, D., Nystrom, A., Zhang, C., Eck, D., Callison-Burch, C., and Carlini, N · 2022
Cited alongside, same era.
ChatGPT: Optimizing Language Models for Dialogue, 2022
OpenAI · 2022
Cited alongside, same era.
Kudugunta, S., Caswell, I., Zhang, B., Garcia, X., Choquette-Choo, C. A., Lee, K., Xin, D., Kusupati, A., Stella, R., Bapna, A., et al · 2023
Closest in time.
Talkin’ ’Bout AI Generation: Copyright and the Generative-AI Supply Chain, 2023
Lee, K., Cooper, A. F., and Grimmelmann, J · 2023
Closest in time.
AI and Law: The Next Generation, 2023
Lee, K., Cooper, A. F., Grimmelmann, J., and Ippolito, D · 2023
Closest in time.
Scaling data-constrained language models
Muennighoff, N., Rush, A. M., Barak, B., Scao, T. L., Piktus, A., Tazi, N., Pyysalo, S., Wolf, T., and Raffel, C · 2023
Closest in time.
Custom instructions for ChatGPT, 2023
OpenAI · 2023
Closest in time.
GPT-4 System Card
OpenAI · 2023
Closest in time.
OpenAI · 2023
Closest in time.
The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only, 2023
Penedo, G., Malartic, Q., Hesslow, D., Cojocaru, R., Cappelli, A., Alobeidli, H., Pannier, B., Almazrouei, E., and Launay, J · 2023
Closest in time.
AI2 Dolma: 3 trillion token open corpus for language model pretraining, 2023
Soldaini, L · 2023
Closest in time.
Diffusion art or digital forgery? Investigating data replication in diffusion models
Somepalli, G., Singla, V., Goldblum, M., Geiping, J., and Goldstein, T · 2023
Closest in time.
LLaMA: Open and Efficient Foundation Language Models, 2023
Touvron, H., Lavril, T., Izacard, G., Martinet, X., Lachaux, M.-A., Lacroix, T., Rozière, B., Goyal, N., Hambro, E., Azhar, F., Rodriguez, A., Joulin, A., Grave, E., and Lample, G · 2023
Closest in time.
LLaMA 2: Open foundation and fine-tuned chat models
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Closest in time.
Universal and transferable adversarial attacks on aligned language models
Zou, A., Wang, Z., Kolter, J. Z., and Fredrikson, M · 2023
Closest in time.