Fetching the paper…
Reading the bibliography…
Questions of fair use of copyright-protected content to train Large Language Models (LLMs) are being actively debated.
Compressive transformers for long-range sequence modelling
Rae, J. W., Potapenko, A., Jayakumar, S. M., Hillier, C., and Lillicrap, T. P · 1911
Earlier work this paper cites.
Membership inference attacks from first principles
Carlini, N., Chien, S., Nasr, M., Song, S., Terzis, A., and Tramer, F · 1914
Earlier work this paper cites.
Project gutenberg, 1971
Hart, M · 1971
Earlier work this paper cites.
Not a word
Alford, H · 2005
Earlier work this paper cites.
Resolving individuals contributing trace amounts of dna to highly complex mixtures using high-density snp genotyping microarrays
Homer, N., Szelinger, S., Redman, M., Duggan, D., Tembe, W., Muehling, J., Pearson, J. V., Stephan, D. A., Nelson, S. F., and Craig, D. W · 2008
Earlier work this paper cites.
Knock knock, who’s there? membership inference on aggregate location data
Pyrgelis, A., Troncoso, C., and De Cristofaro, E · 2017
Earlier work this paper cites.
Membership inference attacks against machine learning models
Shokri, R., Stronati, M., Song, C., and Shmatikov, V · 2017
Earlier work this paper cites.
Ethical challenges in data-driven dialogue systems
Henderson, P., Sinha, K., Angelard-Gontier, N., Ke, N. R., Fried, G., Lowe, R., and Pineau, J · 2018
Earlier work this paper cites.
Comprehensive privacy analysis of deep learning
Nasr, M., Shokri, R., and Houmansadr, A · 2018
Earlier work this paper cites.
Privacy risk in machine learning: Analyzing the connection to overfitting
Yeom, S., Giacomelli, I., Fredrikson, M., and Jha, S · 2018
Earlier work this paper cites.
The secret sharer: Evaluating and testing unintended memorization in neural networks
Carlini, N., Liu, C., Erlingsson, Ú., Kos, J., and Song, D · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners
Radford, A., Wu, J., Child, R., Luan, D., Amodei, D., Sutskever, I., et al · 2019
Earlier work this paper cites.
White-box vs black-box: Bayes optimal strategies for membership inference
Sablayrolles, A., Douze, M., Schmid, C., Ollivier, Y., and Jégou, H · 2019
Earlier work this paper cites.
Auditing data provenance in text-generation models
Song, C. and Shmatikov, V · 2019
Earlier work this paper cites.
Ccnet: Extracting high quality monolingual datasets from web crawl data
Wenzek, G., Lachaux, M.-A., Conneau, A., Chaudhary, V., Guzmán, F., Joulin, A., and Grave, E · 2019
Earlier work this paper cites.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al · 2020
Earlier work this paper cites.
Does learning require memorization? a short tale about a long tail
Feldman, V · 2020
Earlier work this paper cites.
The Pile: An 800gb dataset of diverse text for language modeling
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., Presser, S., and Leahy, C · 2020
Earlier work this paper cites.
Membership inference attacks on sequence-to-sequence models: Is my data in your machine translation system?
Hisamoto, S., Post, M., and Duh, K · 2020
Earlier work this paper cites.
Gutenberg scraper
Pully, K · 2020
Cited alongside, same era.
How much knowledge can you pack into the parameters of a language model?
Roberts, A., Raffel, C., and Shazeer, N · 2020
Cited alongside, same era.
Understanding unintended memorization in federated learning
Thakkar, O., Ramaswamy, S., Mathews, R., and Beaufays, F · 2020
Cited alongside, same era.
Investigating the impact of pre-trained word embeddings on memorization in neural networks
Thomas, A., Adelani, D. I., Davody, A., Mogadala, A., and Klakow, D · 2020
Cited alongside, same era.
Extracting training data from large language models
Carlini, N., Tramer, F., Wallace, E., Jagielski, M., Herbert-Voss, A., Lee, K., Roberts, A., Brown, T., Song, D., Erlingsson, U., et al · 2021
Cited alongside, same era.
Label-only membership inference attacks
Mope: Model perturbation-based privacy attacks on language models
Li, M., Wang, J., Wang, J., and Neel, S · 2023
Later among the works it cites.
Kadrey, silverman, golden v meta platforms, inc
LLMLitigation · 2023
Later among the works it cites.
Membership inference attacks against language models via neighbourhood comparison
Mattern, J., Mireshghallah, F., Jin, Z., Schölkopf, B., Sachan, M., and Berg-Kirkpatrick, T · 2023
Later among the works it cites.
Scaling data-constrained language models, 2023
Muennighoff, N., Rush, A. M., Barak, B., Scao, T. L., Piktus, A., Tazi, N., Pyysalo, S., Wolf, T., and Raffel, C · 2023
Later among the works it cites.
Scalable extraction of training data from (production) language models
Nasr, M., Carlini, N., Hayase, J., Jagielski, M., Cooper, A. F., Ippolito, D., Choquette-Choo, C. A., Wallace, E., Tramèr, F., and Lee, K · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Choquette-Choo, C. A., Tramer, F., Carlini, N., and Papernot, N · 2021
Cited alongside, same era.
Training compute-optimal large language models, 2022
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de Las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Rae, J. W., Vinyals, O., and Sifre, L · 2022
Cited alongside, same era.
Deduplicating training data mitigates privacy risks in language models
Kandpal, N., Wallace, E., and Raffel, C · 2022
Cited alongside, same era.
Deduplicating training data makes language models better
Lee, K., Ippolito, D., Nystrom, A., Zhang, C., Eck, D., Callison-Burch, C., and Carlini, N · 2022
Cited alongside, same era.
Quantifying privacy risks of masked language models using membership inference attacks
Mireshghallah, F., Goyal, K., Uniyal, A., Berg-Kirkpatrick, T., and Shokri, R · 2022
Cited alongside, same era.
Bloom: A 176b-parameter open-access multilingual language model
Scao, T. L., Fan, A., Akiki, C., Pavlick, E., Ilić, S., Hesslow, D., Castagné, R., Luccioni, A. S., Yvon, F., Gallé, M., et al · 2022
Cited alongside, same era.
Opt: Open pre-trained transformer language models, 2022
Zhang, S., Roller, S., Goyal, N., Artetxe, M., Chen, M., Chen, S., Dewan, C., Diab, M., Li, X., Lin, X. V., Mihaylov, T., Ott, M., Shleifer, S., Shuster, K., Simig, D., Koura, P. S., Sridhar, A., Wang, T., and Zettlemoyer, L · 2022
Cited alongside, same era.
Later among the works it cites.
The times sues openai and microsoft over a.i. use of copyrighted work
NewYorkTimes · 2023
Later among the works it cites.
Gpt-4 technical report
OpenAI · 2023
Later among the works it cites.
Penedo, G., Malartic, Q., Hesslow, D., Cojocaru, R., Cappelli, A., Alobeidli, H., Pannier, B., Almazrouei, E., and Launay, J · 2023
Later among the works it cites.
These 183,000 books are fueling the biggest fight in publishing and tech
Reisner, A · 2023
Later among the works it cites.
Generative ai meets copyright
Samuelson, P · 2023
Later among the works it cites.
Detecting pretraining data from large language models
Shi, W., Ajith, A., Xia, M., Huang, Y., Liu, D., Blevins, T., Chen, D., and Zettlemoyer, L · 2023
Later among the works it cites.
SlimPajama: A 627B token cleaned and deduplicated version of RedPajama, June 2023
Soboleva, D., Al-Khateeb, F., Myers, R., Steeves, J. R., Hestness, J., and Dey, N · 2023
Later among the works it cites.
Llama 2: Open foundation and fine-tuned chat models, 2023
Touvron, H., Martin, L., Stone, K., Albert, P., Almahairi, A., Babaei, Y., Bashlykov, N., Batra, S., Bhargava, P., Bhosale, S., et al · 2023
Later among the works it cites.
More than 15,000 authors sign authors guild letter calling on ai industry leaders to protect writers
USAuthorsGuild · 2023
Later among the works it cites.
Common crawl
CommonCrawl · 2024
Closest in time.
Croissantllm: A truly bilingual french-english language model
Faysse, M., Fernandes, P., Guerreiro, N., Loison, A., Alves, D., Corro, C., Boizard, N., Alves, J., Rei, R., Martins, P., et al · 2024
Closest in time.
Jiang, A. Q., Sablayrolles, A., Roux, A., Mensch, A., Savary, B., Bamford, C., Chaplot, D. S., Casas, D. d. l., Hanna, E. B., Bressand, F., et al · 2024
Closest in time.
Madlad-400: A multilingual and document-level large audited dataset
Kudugunta, S., Caswell, I., Zhang, B., Garcia, X., Xin, D., Kusupati, A., Stella, R., Bapna, A., and Firat, O · 2024
Closest in time.
Tinyllama: An open-source small language model, 2024
Zhang, P., Zeng, G., Wang, T., and Lu, W · 2024
Closest in time.