Fetching the paper…
Reading the bibliography…
Practically all large language models have been pre-trained on data that is subject to global uncertainty related to copyright infringement and breach of contract.
44 U.S.C. § 1911 - Free use of Government publications in depositories, 2021c
United States Congress · 1911
Earlier work this paper cites.
Neural networks and physical systems with emergent collective computational abilities
JJ Hopfield · 1982
Earlier work this paper cites.
Learning representations by back-propagating errors
David E Rumelhart, Geoffrey E Hinton, and Ronald J Williams · 1986
Earlier work this paper cites.
The Nature of Copyright: A Law of Users Rights
Lyman Ray Patterson · 1991
Earlier work this paper cites.
How the Internet was indexed
George McMurdo · 1995
Earlier work this paper cites.
Intellectual property and the digital economy: Why the anti-circumvention regulations need to be revised
Pamela Samuelson · 1999
Earlier work this paper cites.
After Napster
Corey Rayburn · 2001
Earlier work this paper cites.
The American National Corpus: More Than the Web Can Provide
Nancy Ide, Randi Reppen, and Keith Suderman · 2002
Earlier work this paper cites.
The American National Corpus: Then, now, and tomorrow
Nancy Ide · 2008
Earlier work this paper cites.
ImageNet: A large-scale hierarchical image database
Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei · 2009
Earlier work this paper cites.
Google Book Search and the future of books in cyberspace
Pamela Samuelson · 2009
Earlier work this paper cites.
Authorship in Wikipedia–Legal Requirements, Community Opinions, and Technical Boundaries
Thomas Roessing · 2010
Earlier work this paper cites.
The copyright principles project: Directions for reform
Pamela Samuelson, Jon A Baumgarten, Michael W Carroll, Julie E Cohen, Troy Dow, Brian Fitzgerald, Laura Gasaway, Daniel Gervais, Terry Ilardi, Jessica Litman, et al · 2010
Earlier work this paper cites.
Commission Decision of 12 December 2011 on the reuse of Commission documents, 2011
European Commission · 2011
Earlier work this paper cites.
Fair Use/Fair Dealing Handbook
Jonathan Band and Jonathan Gerafi · 2013
Earlier work this paper cites.
Dirt Cheap Web-Scale Parallel Text from the Common Crawl
Jason Smith, Herve Saint-Amand, Magdalena Plamadă, Philipp Koehn, Chris Callison-Burch, and Adam Lopez · 2013
Earlier work this paper cites.
Who and what links to the Internet Archive
Yasmin AlNoamany, Ahmed AlSum, Michele C Weigle, and Michael L Nelson · 2014
Earlier work this paper cites.
N-gram counts and language models from the Common Crawl
Christian Buck, Kenneth Heafield, and Bas Van Ooyen · 2014
Earlier work this paper cites.
Free and open source software (FOSS) and other alternative license models: a comparative analysis , volume 12
Axel Metzger · 2015
Earlier work this paper cites.
C4Corpus: Multilingual Web-size corpus with free license
Ivan Habernal, Omnia Zayed, and Iryna Gurevych · 2016
Earlier work this paper cites.
SNAP: A general-purpose network analysis and graph-mining library
Jure Leskovec and Rok Sosič · 2016
Earlier work this paper cites.
CommonCOW: Massively huge web corpora from CommonCrawl data and a method to distribute them freely under restrictive EU copyright laws
Roland Schäfer · 2016
Earlier work this paper cites.
The entropy of words—Learnability and expressivity across more than 1000 languages
Christian Bentz, Dimitrios Alikaniotis, Michael Cysouw, and Ramon Ferrer-i Cancho · 2017
Cited alongside, same era.
Measuring and modeling the US regulatory ecosystem
Michael J Bommarito II and Daniel Martin Katz · 2017
Cited alongside, same era.
Copyright Survives: Rethinking the Copyright-Contract Conflict
Guy A Rub · 2017
Cited alongside, same era.
Compendium of U.S. Copyright Office Practices, 2017
U.S. Copyright Office · 2017
Cited alongside, same era.
Data producer’s right and the protection of machine-generated data
Peter K Yu · 2018
Cited alongside, same era.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al · 2020
Back to the Future: Navigating the Copyright/Contract Interface in the Generate AI Era
Niva Elkin-Koren · 2024
Later among the works it cites.
Differentially Private Next-Token Prediction of Large Language Models
James Flemings, Meisam Razaviyayn, and Murali Annavaram · 2024
Later among the works it cites.
Free Law Project, 2024
Free Law Project · 2024
Later among the works it cites.
Black’s Law Dictionary , volume 12
Bryan A Garner · 2024
Later among the works it cites.
CPR: Retrieval augmented generation for copyright protection
Aditya Golatkar, Alessandro Achille, Luca Zancato, Yu-Xiang Wang, Ashwin Swaminathan, and Stefano Soatto · 2024
Later among the works it cites.
GPT-4 passes the Bar Exam
Daniel Martin Katz, Michael James Bommarito, Shang Gao, and Pablo Arredondo · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The Pile: An 800GB dataset of diverse text for language modeling
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al · 2020
Cited alongside, same era.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J Liu · 2020
Cited alongside, same era.
Georgia v. Public.Resource.Org, Inc., 590 U.S. 255, 2020
Supreme Court of the United States · 2020
Cited alongside, same era.
Measuring law over time: a network analytical framework with an application to statutes and regulations in the United States and Germany
Corinna Coupette, Janis Beckedorf, Dirk Hartung, Michael Bommarito, and Daniel Martin Katz · 2021
Cited alongside, same era.
Pile of Law: Learning responsible data filtering from the law and a 256gb open-source legal dataset
Peter Henderson, Mark Krass, Lucia Zheng, Neel Guha, Christopher D Manning, Dan Jurafsky, and Daniel Ho · 2022
Cited alongside, same era.
Comparing ChatGPT and GPT-4 performance in USMLE soft skill assessments
Dana Brin, Vera Sorin, Akhil Vaid, Ali Soroush, Benjamin S Glicksberg, Alexander W Charney, Girish Nadkarni, and Eyal Klang · 2023
Cited alongside, same era.
Rethinking open source generative AI: open washing and the EU AI Act
Andreas Liesenfeld and Mark Dingemanse · 2024
Later among the works it cites.
A Pretrainer’s Guide to Training Data: Measuring the Effects of Data Age, Domain Coverage, Quality, & Toxicity
Shayne Longpre, Gregory Yauney, Emily Reif, Katherine Lee, Adam Roberts, Barret Zoph, Denny Zhou, Jason Wei, Kevin Robinson, David Mimno, et al · 2024
Later among the works it cites.
SILO language models: Isolating legal risk in a nonparametric datastore
Sewon Min, Suchin Gururangan, Eric Wallace, Hannaneh Hajishirzi, Noah A. Smith, and Luke Zettlemoyer · 2024
Later among the works it cites.
Fairness and Fair Use in Generative AI
Matthew Sag · 2024
Later among the works it cites.
Dolma: an open corpus of three trillion tokens for language model pretraining research
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Evan Walsh, Luke Zettlemoyer, Noah Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, and Kyle Lo · 2024
Later among the works it cites.
RedPajama: an Open Dataset for Training Large Language Models
Maurice Weber, Dan Fu, Quentin Anthony, Yonatan Oren, Shane Adams, Anton Alexandrov, Xiaozhong Lyu, Huu Nguyen, Xiaozhe Yao, Virginia Adams, et al · 2024
Later among the works it cites.
Wikimedia Says No to LLM Training: Internal Communications Reveal Opposition to AI Use, 2024
Jillian Bommarito · 2025
Closest in time.
Michael Bommarito, Daniel Katz, and Jillian Bommarito · 2025
Closest in time.
GPT-4 assistance for improvement of physician performance on patient care tasks: a randomized controlled trial
Ethan Goh, Robert J Gallo, Eric Strong, Yingjie Weng, Hannah Kerman, Jason A Freed, Joséphine A Cool, Zahir Kanjee, Kathleen P Lane, Andrew S Parsons, et al · 2025
Closest in time.
DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al · 2025
Closest in time.
Developers’ Access to Training Data
Adam Jaffe · 2025
Closest in time.
Response to White House Office of Science and Technology (OSTP) for the upcoming US AI Action Plan
OpenAI · 2025
Closest in time.
Open Justice Licence, 2023
The National Archives · 2025
Closest in time.
Terms of Use for USPTO Websites, 2024
United States Patent and Trademark Office · 2025
Closest in time.
CompLex: Legal systems through the lens of complexity science
Pierpaolo Vivo, Daniel M Katz, and JB Ruhl · 2025
Closest in time.