Fetching the paper…
Reading the bibliography…
The proliferation of large language models (LLMs) in the real world has come with a rise in copyright cases against companies for training their models on unlicensed data from the internet.
400: A method for combining non-independent, one-sided tests of significance
Morton B Brown · 1975
Earlier work this paper cites.
Random variables with maximum sums
Ludger Rüschendorf · 1982
Earlier work this paper cites.
Posterior predictive p p -values
Xiao-Li Meng · 1994
Earlier work this paper cites.
Combining dependent p-values
James T Kost and Michael P McDermott · 2002
Earlier work this paper cites.
zlib compression library
Jean-loup Gailly and Mark Adler · 2004
Earlier work this paper cites.
Membership inference attacks against machine learning models
R. Shokri, M. Stronati, C. Song, and V. Shmatikov · 2017
Earlier work this paper cites.
Privacy risk in machine learning: Analyzing the connection to overfitting
S. Yeom, I. Giacomelli, M. Fredrikson, and S. Jha · 2018
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al · 2020
Earlier work this paper cites.
Combining p-values via averaging
Vladimir Vovk and Ruodu Wang · 2020
Earlier work this paper cites.
Extracting training data from large language models
Nicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Úlfar Erlingsson, Alina Oprea, and Colin Raffel · 2021
Earlier work this paper cites.
Dataset inference: Ownership resolution in machine learning
Pratyush Maini, Mohammad Yaghini, and Nicolas Papernot · 2021
Earlier work this paper cites.
Membership inference attacks are easier on difficult problems
Avital Shafran, Shmuel Peleg, and Yedid Hoshen · 2021
Cited alongside, same era.
Membership inference attacks from first principles
Nicholas Carlini, Steve Chien, Milad Nasr, Shuang Song, Andreas Terzis, and Florian Tramèr · 2022
Cited alongside, same era.
Dataset inference for self-supervised models
Adam Dziedzic, Haonan Duan, Muhammad Ahmad Kaleem, Nikita Dhawan, Jonas Guan, Yannis Cattan, Franziska Boenisch, and Nicolas Papernot · 2022
Cited alongside, same era.
Deduplicating training data makes language models better
Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini · 2022
Cited alongside, same era.
URL https://www.bakerlaw.com/getty-images-v-stability-ai/
Getty images vs. stability ai: A landmark case in copyright and ai, 2023 · 2023
Cited alongside, same era.
Pythia: a suite for analyzing large language models across training and scaling
Membership inference attacks against language models via neighbourhood comparison
Justus Mattern, Fatemehsadat Mireshghallah, Zhijing Jin, Bernhard Schoelkopf, Mrinmaya Sachan, and Taylor Berg-Kirkpatrick · 2023
Later among the works it cites.
Silo language models: Isolating legal risk in a nonparametric datastore
Sewon Min, Suchin Gururangan, Eric Wallace, Hannaneh Hajishirzi, Noah A Smith, and Luke Zettlemoyer · 2023
Later among the works it cites.
Detectgpt: zero-shot machine-generated text detection using probability curvature
Eric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning, and Chelsea Finn · 2023
Later among the works it cites.
Beyond fair use: Legal risk evaluation for training llms on copyrighted text
Noorjahan Rahman and Eduardo Santacana · 2023
Later among the works it cites.
Privacy auditing with one (1) training run
Thomas Steinke, Milad Nasr, and Matthew Jagielski · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Stella Biderman, Hailey Schoelkopf, Quentin Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, Aviya Skowron, Lintang Sutawika, and Oskar Van Der Wal · 2023
Cited alongside, same era.
Nl-augmenter: A framework for task-sensitive natural language augmentation, 2023
Kaustubh D. Dhole, Varun Gangal, Sebastian Gehrmann, Aadesh Gupta, Zhenhao Li, Saad Mahamood, Abinaya Mahendiran, Simon Mille, Ashish Srivastava, Samson Tan, Tongshuang Wu, Jascha Sohl-Dickstein, Jinho D. Choi, Eduard Hovy, Ondrej Dusek, Sebastian Ruder, Sajant Anand, Nagender Aneja, Rabin Banjade, Lisa Barthe, Hanna Behnke, Ian Berlot-Attwell, Connor Boyle, Caroline Brun, Marco Antonio Sobrevilla Cabezudo, Samuel Cahyawijaya, Emile Chapuis, Wanxiang Che, Mukund Choudhary, Christian Clauss, Pierre Colombo, Filip Cornell, Gautier Dagan, Mayukh Das, Tanay Dixit, Thomas Dopierre, Paul-Alexis Dray, Suchitra Dubey, Tatiana Ekeinhor, Marco Di Giovanni, Rishabh Gupta, Rishabh Gupta, Louanes Hamla, Sang Han, Fabrice Harel-Canada, Antoine Honore, Ishan Jindal, Przemyslaw K. Joniak, Denis Kleyko, Venelin Kovatchev, Kalpesh Krishna, Ashutosh Kumar, Stefan Langer, Seungjae Ryan Lee, Corey James Levinson, Hualou Liang, Kaizhao Liang, Zhexiong Liu, Andrey Lukyanenko, Vukosi Marivate, Gerard de Melo, Simon Meoni, Maxime Meyer, Afnan Mir, Nafise Sadat Moosavi, Niklas Muennighoff, Timothy Sum Hon Mun, Kenton Murray, Marcin Namysl, Maria Obedkova, Priti Oli, Nivranshu Pasricha, Jan Pfister, Richard Plant, Vinay Prabhu, Vasile Pais, Libo Qin, Shahab Raji, Pawan Kumar Rajpoot, Vikas Raunak, Roy Rinberg, Nicolas Roberts, Juan Diego Rodriguez, Claude Roux, Vasconcellos P. H. S., Ananya B. Sai, Robin M. Schmidt, Thomas Scialom, Tshephisho Sefara, Saqib N. Shamsi, Xudong Shen, Haoyue Shi, Yiwen Shi, Anna Shvets, Nick Siegel, Damien Sileo, Jamie Simon, Chandan Singh, Roman Sitelew, Priyank Soni, Taylor Sorensen, William Soto, Aman Srivastava, KV Aditya Srivatsa, Tony Sun, Mukund Varma T, A Tabassum, Fiona Anting Tan, Ryan Teehan, Mo Tiwari, Marie Tolkiehn, Athena Wang, Zijian Wang, Gloria Wang, Zijie J. Wang, Fuxuan Wei, Bryan Wilie, Genta Indra Winata, Xinyi Wu, Witold Wydmański, Tianbao Xie, Usama Yaseen, M. Yee, Jing Zhang, and Yue Zhang · 2023
Cited alongside, same era.
Flocks of stochastic parrots: Differentially private prompt learning for large language models
Haonan Duan, Adam Dziedzic, Nicolas Papernot, and Franziska Boenisch · 2023
Cited alongside, same era.
Tinystories: How small can language models be and still speak coherent english?
Ronen Eldan and Yuanzhi Li · 2023
Cited alongside, same era.
Textbooks are all you need ii: phi-1.5 technical report
Yuanzhi Li, Sébastien Bubeck, Ronen Eldan, Allie Del Giorno, Suriya Gunasekar, and Yin Tat Lee · 2023
Cited alongside, same era.
URL https://gemini.google.com/
Germini, https://gemini.google.com/
Cited in the paper.
URL https://ai.meta.com/blog/meta-llama-3/
Introducing meta llama 3: The most capable openly available llm to date, https://ai.meta.com/blog/meta-llama-3/
Cited in the paper.
Unveiling security, privacy, and ethical concerns of chatgpt
Xiaodong Wu, Ran Duan, and Jianbing Ni · 2023
Later among the works it cites.
Fast-detectGPT: Efficient zero-shot detection of machine-generated text via conditional probability curvature
Guangsheng Bao, Yanbin Zhao, Zhiyang Teng, Linyi Yang, and Yue Zhang · 2024
Closest in time.
Do membership inference attacks work on large language models?
Michael Duan, Anshuman Suri, Niloofar Mireshghallah, Sewon Min, Weijia Shi, Luke Zettlemoyer, Yulia Tsvetkov, Yejin Choi, David Evans, and Hannaneh Hajishirzi · 2024
Closest in time.
Proving test set contamination for black-box language models
Yonatan Oren, Nicole Meister, Niladri S. Chatterji, Faisal Ladhak, and Tatsunori Hashimoto · 2024
Closest in time.
Detecting pretraining data from large language models
Weijia Shi, Anirudh Ajith, Mengzhou Xia, Yangsibo Huang, Daogao Liu, Terra Blevins, Danqi Chen, and Luke Zettlemoyer · 2024
Closest in time.