Fetching the paper…
Reading the bibliography…
Large language models (LLMs) are typically trained on enormous quantities of unlicensed text, a practice that has led to scrutiny due to possible intellectual property infringement and ethical concerns.
Socialiqa: Commonsense reasoning about social interactions, 2019
Maarten Sap, Hannah Rashkin, Derek Chen, Ronan LeBras, and Yejin Choi · 1904
Earlier work this paper cites.
Boolq: Exploring the surprising difficulty of natural yes/no questions
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova · 1905
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language, 2019
Yonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao, and Yejin Choi · 1911
Earlier work this paper cites.
Compressive transformers for long-range sequence modelling
Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, Chloe Hillier, and Timothy P Lillicrap · 1911
Earlier work this paper cites.
Subject matter of copyright: In general, December 1990
17 U.S. Code § 102 · 1990
Earlier work this paper cites.
A new algorithm for data compression
Philip Gage · 1994
Earlier work this paper cites.
Learning question classifiers
Xin Li and D. Roth · 2002
Earlier work this paper cites.
Logic2text: High-fidelity natural language generation from logical forms
Zhiyu Chen, Wenhu Chen, Hanwen Zha, Xiyou Zhou, Yunkai Zhang, Sairam Sundaresan, and William Yang Wang · 2004
Earlier work this paper cites.
The winograd schema challenge
H. Levesque, E. Davis, and L. Morgenstern · 2011
Earlier work this paper cites.
Large text compression benchmark, 2011
Matt Mahoney · 2011
Earlier work this paper cites.
Quantitative analysis of culture using millions of digitized books
Jean-Baptiste Michel, Yuan Kui Shen, Aviva Presser Aiden, Adrian Veres, Matthew K. Gray, The Google Books Team, Joseph P. Pickett, Dale Hoiberg, Dan Clancy, Peter Norvig, Jon Orwant, Steven Pinker, Martin A. Nowak, and Erez Lieberman Aiden · 2011
Earlier work this paper cites.
Open-domain anatomical entity mention detection
Tomoko Ohta, Sampo Pyysalo, Junichi Tsujii, and S. Ananiadou · 2012
Earlier work this paper cites.
Semantic parsing on freebase from question-answer pairs
Jonathan Berant, A. Chou, Roy Frostig, and Percy Liang · 2013
Earlier work this paper cites.
Using the framework method for the analysis of qualitative data in multi-disciplinary health research
N. Gale, G. Heath, E. Cameron, S. Rashid, and S. Redwood · 2013
Earlier work this paper cites.
Dbpedia - a large-scale, multilingual knowledge base extracted from wikipedia
Jens Lehmann, Robert Isele, Max Jakob, Anja Jentzsch, D. Kontokostas, Pablo N. Mendes, Sebastian Hellmann, M. Morsey, Patrick van Kleef, S. Auer, and Christian Bizer · 2015
Earlier work this paper cites.
Librispeech: An asr corpus based on public domain audio books
Vassil Panayotov, Guoguo Chen, Daniel Povey, and S. Khudanpur · 2015
Earlier work this paper cites.
C4Corpus: Multilingual web-size corpus with free license
Ivan Habernal, Omnia Zayed, and Iryna Gurevych · 2016
Earlier work this paper cites.
Neural text generation from structured data with application to the biography domain
R. Lebret, David Grangier, and Michael Auli · 2016
Earlier work this paper cites.
Pointer sentinel mixture models
Stephen Merity, Caiming Xiong, James Bradbury, and R. Socher · 2016
Earlier work this paper cites.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Earlier work this paper cites.
Introduction of the asian language treebank
Hammam Riza, Michael Purwoadi, Gunarso, Teduh Uliniansyah, Aw Ai Ti, Sharifah Mahani Aljunied, Luong Chi Mai, V. Thang, N. Thai, Vichet Chea, Rapid Sun, Sethserey Sam, Sopheap Seng, K. Soe, K. Nwet, M. Utiyama, and Chenchen Ding · 2016
Earlier work this paper cites.
Toxic comment classification challenge
cjadams, Jeffrey Sorensen, Julia Elliott, Lucas Dixon, Mark McDonald, nithum, and Will Cukierski · 2017
Earlier work this paper cites.
Harvesting common-sense navigational knowledge for robotics from uncurated text corpora
Nancy Fulda, Nathan Tibbetts, Zachary Brown, and D. Wingate · 2017
Earlier work this paper cites.
Bag of tricks for efficient text classification
Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov · 2017
Earlier work this paper cites.
The e2e dataset: New challenges for end-to-end generation
Jekaterina Novikova, Ondrej Dusek, and Verena Rieser · 2017
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin · 2017
Earlier work this paper cites.
Elastic ChatNoir: Search Engine for the ClueWeb and the Common Crawl
Janek Bevendorff, Benno Stein, Matthias Hagen, and Martin Potthast · 2018
Earlier work this paper cites.
e-snli: Natural language inference with natural language explanations
Oana-Maria Camburu, Tim Rocktäschel, Thomas Lukasiewicz, and Phil Blunsom · 2018
Earlier work this paper cites.
Quac: Question answering in context
Eunsol Choi, He He, Mohit Iyyer, Mark Yatskar, Wen tau Yih, Yejin Choi, Percy Liang, and Luke Zettlemoyer · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord · 2018
Earlier work this paper cites.
Tracking state changes in procedural text: a challenge dataset and models for process paragraph comprehension
Bhavana Dalvi, Lifu Huang, Niket Tandon, Wen tau Yih, and Peter Clark · 2018
Earlier work this paper cites.
Breaking nli systems with sentences that require simple lexical inferences
Max Glockner, Vered Shwartz, and Yoav Goldberg · 2018
Earlier work this paper cites.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder · 2018
Earlier work this paper cites.
Scitail: A textual entailment dataset from science question answering
Tushar Khot, Ashish Sabharwal, and Peter Clark · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training, 2018
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever · 2018
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for squad
Pranav Rajpurkar, Robin Jia, and Percy Liang · 2018
Earlier work this paper cites.
Commonsenseqa: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant · 2018
Earlier work this paper cites.
Fever: a large-scale dataset for fact extraction and verification
James Thorne, Andreas Vlachos, Christos Christodoulopoulos, and Arpit Mittal · 2018
Earlier work this paper cites.
Airdialogue: An environment for goal-oriented dialogue research
Wei Wei, Quoc V. Le, Andrew M. Dai, and Jia Li · 2018
Earlier work this paper cites.
Hotpotqa: A dataset for diverse, explainable multi-hop question answering
Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, R. Salakhutdinov, and Christopher D. Manning · 2018
Earlier work this paper cites.
On the cross-lingual transferability of monolingual representations
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama · 2019
Earlier work this paper cites.
What does BERT look at? an analysis of BERT’s attention
Kevin Clark, Urvashi Khandelwal, Omer Levy, and Christopher D Manning · 2019
Earlier work this paper cites.
A span-extraction dataset for chinese machine reading comprehension
Yiming Cui, Ting Liu, Li Xiao, Zhipeng Chen, Wentao Ma, Wanxiang Che, Shijin Wang, and Guoping Hu · 2019
Earlier work this paper cites.
Quoref: A reading comprehension dataset with questions requiring coreferential reasoning
Pradeep Dasigi, Nelson F. Liu, Ana Marasović, Noah A. Smith, and Matt Gardner · 2019
Earlier work this paper cites.
Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner · 2019
Earlier work this paper cites.
Where’s my head? definition, data set, and models for numeric fused-head identification and resolution
Yanai Elazar and Yoav Goldberg · 2019
Earlier work this paper cites.
When choosing plausible alternatives, clever hans can be clever
Pride Kavumba, Naoya Inoue, Benjamin Heinzerling, Keshav Singh, Paul Reisert, and Kentaro Inui · 2019
Earlier work this paper cites.
Natural questions: A benchmark for question answering research
T. Kwiatkowski, J. Palomaki, Olivia Redfield, Michael Collins, Ankur P. Parikh, Chris Alberti, D. Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc V. Le, and Slav Petrov · 2019
Earlier work this paper cites.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2019
Earlier work this paper cites.
Directive (eu) 2019/790, 2019
European Parliament and Council of the European Union · 2019
Earlier work this paper cites.
Language models as knowledge bases?
Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners, 2019
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever · 2019
Earlier work this paper cites.
Coached conversational preference elicitation: A case study in understanding movie preferences
Filip Radlinski, K. Balog, B. Byrne, and K. Krishnamoorthi · 2019
Earlier work this paper cites.
Explain yourself! leveraging language models for commonsense reasoning
Nazneen Rajani, Bryan McCann, Caiming Xiong, and R. Socher · 2019
Earlier work this paper cites.
Winogrande
Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi · 2019
Earlier work this paper cites.
Evaluating gender bias in machine translation
Gabriel Stanovsky, Noah A. Smith, and Luke Zettlemoyer · 2019
Earlier work this paper cites.
Asynchronous pipeline for processing huge corpora on medium to low resource infrastructures
Pedro Javier Ortiz Suárez, Benoît Sagot, and Laurent Romary · 2019
Earlier work this paper cites.
Quartz: An open-domain dataset of qualitative relationship questions
Oyvind Tafjord, Matt Gardner, Kevin Lin, and Peter Clark · 2019
Earlier work this paper cites.
Resolving gendered ambiguous pronouns with bert
Kellie Webster, Marta Recasens, Vera Axelrod, and Jason Baldridge · 2019
Cited alongside, same era.
Hellaswag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi · 2019
Cited alongside, same era.
Beat the ai: Investigating adversarial human annotation for reading comprehension
Max Bartolo, A. Roberts, Johannes Welbl, Sebastian Riedel, and Pontus Stenetorp · 2020
Cited alongside, same era.
HybridQA: A dataset of multi-hop question answering over tabular and textual data
Wenhu Chen, Hanwen Zha, Zhiyu Chen, Wenhan Xiong, Hong Wang, and William Yang Wang · 2020
Cited alongside, same era.
Moral stories: Situated reasoning about norms, intents, actions, and their consequences
Denis Emelin, Ronan Le Bras, Jena D. Hwang, Maxwell Forbes, and Yejin Choi · 2020
Llm360: Towards fully transparent open-source llms, 2023
Zhengzhong Liu, Aurick Qiao, Willie Neiswanger, Hongyi Wang, Bowen Tan, Tianhua Tao, Junbo Li, Yuqi Wang, Suqi Sun, Omkar Pangarkar, Richard Fan, Yi Gu, Victor Miller, Yonghao Zhuang, Guowei He, Haonan Li, Fajri Koto, Liping Tang, Nikhil Ranjan, Zhiqiang Shen, Xuguang Ren, Roberto Iriondo, Cun Mu, Zhiting Hu, Mark Schulze, Preslav Nakov, Tim Baldwin, and Eric P. Xing · 2023
Later among the works it cites.
Discit ergo est: Training data provenance and fair use
Robert Mahari and Shayne Longpre · 2023
Later among the works it cites.
Scaling data-constrained language models
Niklas Muennighoff, Alexander Rush, Boaz Barak, Teven Le Scao, Nouamane Tazi, Aleksandra Piktus, Sampo Pyysalo, Thomas Wolf, and Colin A Raffel · 2023
Later among the works it cites.
Trak: Attributing model behavior at scale, 2023
Sung Min Park, Kristian Georgiev, Andrew Ilyas, Guillaume Leclerc, and Aleksander Madry · 2023
Later among the works it cites.
peS2o (Pretraining Efficiently on S2ORC) Dataset
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Iirc: A dataset of incomplete information reading comprehension questions
James Ferguson, Matt Gardner, Tushar Khot, and Pradeep Dasigi · 2020
Cited alongside, same era.
The pile: An 800gb dataset of diverse text for language modeling, 2020
Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy · 2020
Cited alongside, same era.
Wiki-40b: Multilingual language model dataset
Mandy Guo, Zihang Dai, Denny Vrandečić, and Rami Al-Rfou · 2020
Cited alongside, same era.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt · 2020
Cited alongside, same era.
Comet-atomic 2020: On symbolic and neural commonsense knowledge graphs
Jena D. Hwang, Chandra Bhagavatula, Ronan Le Bras, Jeff Da, Keisuke Sakaguchi, Antoine Bosselut, and Yejin Choi · 2020
Cited alongside, same era.
Wikilingua: A new benchmark dataset for multilingual abstractive summarization
Faisal Ladhak, Esin Durmus, Claire Cardie, and K. McKeown · 2020
Cited alongside, same era.
S2ORC: The semantic scholar open research corpus
Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, and Daniel Weld · 2020
Cited alongside, same era.
Luca Soldaini and Kyle Lo · 2023
Later among the works it cites.
Introducing mpt-7b: A new standard for open-source, commercially usable llms, 2023
MosaicML NLP Team · 2023
Later among the works it cites.
Helpsteer: Multi-attribute helpfulness dataset for steerlm
Zhilin Wang, Yi Dong, Jiaqi Zeng, Virginia Adams, Makesh Narsimhan Sreedhar, Daniel Egert, Olivier Delalleau, Jane Polak Scowcroft, Neel Kant, Aidan Swope, and Oleksii Kuchaiev · 2023
Later among the works it cites.
Doremi: Optimizing data mixtures speeds up language model pretraining
Sang Michael Xie, Hieu Pham, Xuanyi Dong, Nan Du, Hanxiao Liu, Yifeng Lu, Percy S Liang, Quoc V Le, Tengyu Ma, and Adams Wei Yu · 2023
Later among the works it cites.
Subject matter of copyright: United States Government works, December 2024
17 U.S. Code § 105 · 2024
Later among the works it cites.
A survey on data selection for language models
Alon Albalak, Yanai Elazar, Sang Michael Xie, Shayne Longpre, Nathan Lambert, Xinyi Wang, Niklas Muennighoff, Bairu Hou, Liangming Pan, Haewon Jeong, et al · 2024
Later among the works it cites.
To code, or not to code? exploring impact of code in pre-training
Viraat Aryabumi, Yixuan Su, Raymond Ma, Adrien Morisot, Ivan Zhang, Acyr Locatelli, Marzieh Fadaee, Ahmet Üstün, and Sara Hooker · 2024
Later among the works it cites.
Stable lm 2 1.6 b technical report
Marco Bellagente, Jonathan Tow, Dakota Mahan, Duy Phung, Maksym Zhuravinskyi, Reshinth Adithyan, James Baicoianu, Ben Brooks, Nathan Cooper, Ashish Datta, et al · 2024
Later among the works it cites.
Deepseek llm: Scaling open-source language models with longtermism
Xiao Bi, Deli Chen, Guanting Chen, Shanhuang Chen, Damai Dai, Chengqi Deng, Honghui Ding, Kai Dong, Qiushi Du, Zhe Fu, et al · 2024
Later among the works it cites.
URL https://chatgptiseatingtheworld.com
Chat GPT Is Eating the World, 2024 · 2024
Later among the works it cites.
What is your data worth to gpt? llm-scale data valuation with influence functions, 2024
Sang Keun Choe, Hwijeen Ahn, Juhan Bae, Kewen Zhao, Minsoo Kang, Youngseog Chung, Adithya Pratapa, Willie Neiswanger, Emma Strubell, Teruko Mitamura, Jeff Schneider, Eduard Hovy, Roger Grosse, and Eric Xing · 2024
Later among the works it cites.
The Files are in the Computer: Copyright, Memorization, and Generative AI
A. Feder Cooper and James Grimmelmann · 2024
Later among the works it cites.
Do membership inference attacks work on large language models?
Michael Duan, Anshuman Suri, Niloofar Mireshghallah, Sewon Min, Weijia Shi, Luke Zettlemoyer, Yulia Tsvetkov, Yejin Choi, David Evans, and Hannaneh Hajishirzi · 2024
Later among the works it cites.
CommonCanvas: Open Diffusion Models Trained on Creative-Commons Images
Aaron Gokaslan, A. Feder Cooper, Jasmine Collins, Landan Seguin, Austin Jacobson, Mihir Patel, Jonathan Frankle, Cory Stephenson, and Volodymyr Kuleshov · 2024
Later among the works it cites.
Olmo: Accelerating the science of language models, 2024
Dirk Groeneveld, Iz Beltagy, Pete Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Harsh Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, Shane Arora, David Atkinson, Russell Authur, Khyathi Raghavi Chandu, Arman Cohan, Jennifer Dumas, Yanai Elazar, Yuling Gu, Jack Hessel, Tushar Khot, William Merrill, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Valentina Pyatkin, Abhilasha Ravichander, Dustin Schwenk, Saurabh Shah, Will Smith, Emma Strubell, Nishant Subramani, Mitchell Wortsman, Pradeep Dasigi, Nathan Lambert, Kyle Richardson, Luke Zettlemoyer, Jesse Dodge, Kyle Lo, Luca Soldaini, Noah A. Smith, and Hannaneh Hajishirzi · 2024
Later among the works it cites.
AI Training and Copyright Infringement: Solutions from Asia, October 2024
Seth Hays · 2024
Later among the works it cites.
Minicpm: Unveiling the potential of small language models with scalable training strategies
Shengding Hu, Yuge Tu, Xu Han, Chaoqun He, Ganqu Cui, Xiang Long, Zhi Zheng, Yewei Fang, Yuxiang Huang, Weilin Zhao, et al · 2024
Later among the works it cites.
Simple and scalable strategies to continually pre-train large language models
Adam Ibrahim, Benjamin Thérien, Kshitij Gupta, Mats L Richter, Quentin Anthony, Timothée Lesort, Eugene Belilovsky, and Irina Rish · 2024
Later among the works it cites.
Model AI Governance Framework for Generative AI: Fostering a Trusted Ecosystem, May 2024
Infocomm Media Development Authority of Singapore (IMDA), Aicadium, and AI Verify Foundation · 2024
Later among the works it cites.
Releasing Common Corpus: the largest public domain dataset for training LLMs, 2024
Pierre-Carl Langlais · 2024
Later among the works it cites.
AI White Paper 2024: New Strategies in Stage II, Toward the world’s most AI-friendly country, April 2024
LDP Headquarters for the Promotion of Digital Society and Project Team on the Evolution and Implementation of AIs · 2024
Later among the works it cites.
A pretrainer’s guide to training data: Measuring the effects of data age, domain coverage, quality, & toxicity
Shayne Longpre, Gregory Yauney, Emily Reif, Katherine Lee, Adam Roberts, Barret Zoph, Denny Zhou, Jason Wei, Kevin Robinson, David Mimno, et al · 2024
Later among the works it cites.
Starcoder 2 and the stack v2: The next generation, 2024
Anton Lozhkov, Raymond Li, Loubna Ben Allal, Federico Cassano, Joel Lamy-Poirier, Nouamane Tazi, Ao Tang, Dmytro Pykhtar, Jiawei Liu, Yuxiang Wei, Tianyang Liu, Max Tian, Denis Kocetkov, Arthur Zucker, Younes Belkada, Zijian Wang, Qian Liu, Dmitry Abulkhanov, Indraneil Paul, Zhuang Li, Wen-Ding Li, Megan Risdal, Jia Li, Jian Zhu, Terry Yue Zhuo, Evgenii Zheltonozhskii, Nii Osae Osae Dade, Wenhao Yu, Lucas Krauß, Naman Jain, Yixuan Su, Xuanli He, Manan Dey, Edoardo Abati, Yekun Chai, Niklas Muennighoff, Xiangru Tang, Muhtasham Oblokulov, Christopher Akiki, Marc Marone, Chenghao Mou, Mayank Mishra, Alex Gu, Binyuan Hui, Tri Dao, Armel Zebaze, Olivier Dehaene, Nicolas Patry, Canwen Xu, Julian McAuley, Han Hu, Torsten Scholak, Sebastien Paquet, Jennifer Robinson, Carolyn Jane Anderson, Nicolas Chapados, Mostofa Patwary, Nima Tajbakhsh, Yacine Jernite, Carlos Muñoz Ferrandis, Lingming Zhang, Sean Hughes, Thomas Wolf, Arjun Guha, Leandro von Werra, and Harm de Vries · 2024
Later among the works it cites.
SILO language models: Isolating legal risk in a nonparametric datastore
Sewon Min, Suchin Gururangan, Eric Wallace, Weijia Shi, Hannaneh Hajishirzi, Noah A. Smith, and Luke Zettlemoyer · 2024
Later among the works it cites.
The FineWeb datasets: Decanting the web for the finest text data at scale
Guilherme Penedo, Hynek Kydlíček, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, and Thomas Wolf · 2024
Later among the works it cites.
How to Think About Remedies in the Generative AI Copyright Cases
Pamela Samuelson · 2024
Later among the works it cites.
Aya dataset: An open-access collection for multilingual instruction tuning
Shivalika Singh, Freddie Vargus, Daniel Dsouza, B"orje F. Karlsson, Abinaya Mahendiran, Wei-Yin Ko, Herumb Shandilya, Jay Patel, Deividas Mataciunas, Laura OMahony, Mike Zhang, Ramith Hettiarachchi, Joseph Wilson, Marina Machado, Luisa Souza Moura, Dominik Krzemi’nski, Hakimeh Fadaei, Irem Ergun, Ifeoma Okoh, Aisha Alaagib, Oshan Mudannayake, Zaid Alyafeai, Minh Chien Vu, Sebastian Ruder, Surya Guthikonda, Emad A. Alghamdi, Sebastian Gehrmann, Niklas Muennighoff, Max Bartolo, Julia Kreutzer, A. Ustun, Marzieh Fadaee, and Sara Hooker · 2024
Later among the works it cites.
Dolma: an open corpus of three trillion tokens for language model pretraining research
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh, Luke Zettlemoyer, Noah A. Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, and Kyle Lo · 2024
Later among the works it cites.
Txt360: A top-quality llm pre-training dataset requires the perfect blend, 2024
Liping Tang, Nikhil Ranjan, Omkar Pangarkar, Xuezhi Liang, Zhen Wang, Li An, Bhaskar Rao, Linghao Jin, Huijuan Wang, Zhoujun Cheng, Suqi Sun, Cun Mu, Victor Miller, Xuezhe Ma, Yue Peng, Zhengzhong Liu, and Eric P. Xing · 2024
Later among the works it cites.
Meta Lingua: A minimal PyTorch LLM training library, 2024
Mathurin Videau, Badr Youbi Idrissi, Daniel Haziza, Luca Wehrstedt, Jade Copet, Olivier Teytaud, and David Lopez-Paz · 2024
Later among the works it cites.
Redpajama: an open dataset for training large language models, 2024
Maurice Weber, Daniel Fu, Quentin Anthony, Yonatan Oren, Shane Adams, Anton Alexandrov, Xiaozhong Lyu, Huu Nguyen, Xiaozhe Yao, Virginia Adams, Ben Athiwaratkun, Rahul Chalamala, Kezhen Chen, Max Ryabinin, Tri Dao, Percy Liang, Christopher Ré, Irina Rish, and Ce Zhang · 2024
Later among the works it cites.
OpenAI prepares to fight for its life as legal troubles mount
Cat Zakrzewski, Nitasha Tiku, and Elizabeth Dwoskin · 2024
Later among the works it cites.
MAP-Neo: Highly capable and transparent bilingual large language model series
Ge Zhang, Scott Qu, Jiaheng Liu, Chenchen Zhang, Chenghua Lin, Chou Leuang Yu, Danny Pan, Esther Cheng, Jie Liu, Qunshu Lin, Raven Yuan, Tuney Zheng, Wei Pang, Xinrun Du, Yiming Liang, Yinghao Ma, Yizhi Li, Ziyang Ma, Bill Lin, Emmanouil Benetos, Huan Yang, Junting Zhou, Kaijing Ma, Minghao Liu, Morry Niu, Noah Wang, Quehry Que, Ruibo Liu, Sine Liu, Shawn Guo, Soren Gao, Wangchunshu Zhou, Xinyue Zhang, Yizhi Zhou, Yubo Wang, Yuelin Bai, Yuhan Zhang, Yuxiang Zhang, Zenith Wang, Zhenzhu Yang, Zijian Zhao, Jiajun Zhang, Wanli Ouyang, Wenhao Huang, and Wenhu Chen · 2024
Later among the works it cites.
Wildchat: 1m chatGPT interaction logs in the wild
Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng · 2024
Later among the works it cites.
Smollm2: When smol goes big – data-centric training of a small language model, 2025
Loubna Ben Allal, Anton Lozhkov, Elie Bakouch, Gabriel Martín Blázquez, Guilherme Penedo, Lewis Tunstall, Andrés Marafioti, Hynek Kydlíček, Agustín Piqueres Lajarín, Vaibhav Srivastav, Joshua Lochner, Caleb Fahlgren, Xuan-Son Nguyen, Clémentine Fourrier, Ben Burtenshaw, Hugo Larcher, Haojun Zhao, Cyril Zakka, Mathieu Morlon, Colin Raffel, Leandro von Werra, and Thomas Wolf · 2025
Closest in time.
Towards best practices for open datasets for llm training, 2025
Stefan Baack, Stella Biderman, Kasia Odrozek, Aviya Skowron, Ayah Bdeir, Jillian Bommarito, Jennifer Ding, Maximilian Gahntz, Paul Keller, Pierre-Carl Langlais, Greg Lindahl, Sebastian Majstorovic, Nik Marda, Guilherme Penedo, Maarten Van Segbroeck, Jennifer Wang, Leandro von Werra, Mitchell Baker, Julie Belião, Kasia Chmielinski, Marzieh Fadaee, Lisa Gutermuth, Hynek Kydlíček, Greg Leppert, EM Lewis-Jong, Solana Larsen, Shayne Longpre, Angela Oduor Lungati, Cullen Miller, Victor Miller, Max Ryabinin, Kathleen Siminyu, Andrew Strait, Mark Surman, Anna Tumadóttir, Maurice Weber, Rebecca Weiss, Lee White, and Thomas Wolf · 2025
Closest in time.
License List (version 15), 2025
Blue Oak Council · 2025
Closest in time.
Extracting memorized pieces of (copyrighted) books from open-weight language models
A. Feder Cooper, Aaron Gokaslan, Amy B. Cyphert, Christopher De Sa, Mark A. Lemley, Daniel E. Ho, and Percy Liang · 2025
Closest in time.
Datasets, documents, and repetitions: The practicalities of unequal data quality
Alex Fang, Hadi Pouransari, Matt Jordan, Alexander Toshev, Vaishaal Shankar, Ludwig Schmidt, and Tom Gunter · 2025
Closest in time.
Olmes: A standard for language model evaluations, 2025
Yuling Gu, Oyvind Tafjord, Bailey Kuehl, Dany Haddad, Jesse Dodge, and Hannaneh Hajishirzi · 2025
Closest in time.
URL https://huggingface.co/datasets/PleIAs/common_corpus
HuggingFace: Common Corpus, 2025 · 2025
Closest in time.
The kl3m data project: Copyright-clean training resources for large language models, 2025
Michael J Bommarito II, Jillian Bommarito, and Daniel Martin Katz · 2025
Closest in time.
Position: The most expensive part of an llm should be its training data
Nikhil Kandpal and Colin Raffel · 2025
Closest in time.
Datacomp-lm: In search of the next generation of training sets for language models, 2025
Jeffrey Li, Alex Fang, Georgios Smyrnis, Maor Ivgi, Matt Jordan, Samir Gadre, Hritik Bansal, Etash Guha, Sedrick Keh, Kushal Arora, Saurabh Garg, Rui Xin, Niklas Muennighoff, Reinhard Heckel, Jean Mercat, Mayee Chen, Suchin Gururangan, Mitchell Wortsman, Alon Albalak, Yonatan Bitton, Marianna Nezhurina, Amro Abbas, Cheng-Yu Hsieh, Dhruba Ghosh, Josh Gardner, Maciej Kilian, Hanlin Zhang, Rulin Shao, Sarah Pratt, Sunny Sanyal, Gabriel Ilharco, Giannis Daras, Kalyani Marathe, Aaron Gokaslan, Jieyu Zhang, Khyathi Chandu, Thao Nguyen, Igor Vasiljevic, Sham Kakade, Shuran Song, Sujay Sanghavi, Fartash Faghri, Sewoong Oh, Luke Zettlemoyer, Kyle Lo, Alaaeldin El-Nouby, Hadi Pouransari, Alexander Toshev, Stephanie Wang, Dirk Groeneveld, Luca Soldaini, Pang Wei Koh, Jenia Jitsev, Thomas Kollar, Alexandros G. Dimakis, Yair Carmon, Achal Dave, Ludwig Schmidt, and Vaishaal Shankar · 2025
Closest in time.
URL https://openalex.org
OpenAlex, 2025 · 2025
Closest in time.
Privacy auditing of large language models
Ashwinee Panda, Xinyu Tang, Milad Nasr, Christopher A Choquette-Choo, and Prateek Mittal · 2025
Closest in time.
Parser for uk parliament proceedings
ParlParse · 2025
Closest in time.
The size of datasets used to train language models doubles approximately every seven months, 2024
Robi Rahman and David Owen · 2025
Closest in time.
How much is enough? the diminishing returns of tokenization training data, 2025
Varshini Reddy, Craig W. Schmidt, Yuval Pinter, and Chris Tanner · 2025
Closest in time.
The globalization of copyright exceptions for ai training
Matthew Sag and Peter K. Yu · 2025
Closest in time.
Qwen3, April 2025
Qwen Team · 2025
Closest in time.
Mixmin: Finding data mixtures via convex minimization, 2025
Anvith Thudi, Evianne Rovers, Yangjun Ruan, Tristan Thrush, and Chris J. Maddison · 2025
Closest in time.
Open parliament license
UK Parliament · 2025
Closest in time.
Organize the web: Constructing domains enhances pre-training data curation
Alexander Wettig, Kyle Lo, Sewon Min, Hannaneh Hajishirzi, Danqi Chen, and Luca Soldaini · 2025
Closest in time.