Fetching the paper…
Reading the bibliography…
Training a state-of-the-art Large Language Model (LLM) is an increasingly expensive endeavor due to growing computational, hardware, energy, and engineering demands.
Data shapley: Equitable valuation of data for machine learning, 2019
Ghorbani, A. and Zou, J · 1904
Earlier work this paper cites.
Scaling laws for neural language models, 2020
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2001
Earlier work this paper cites.
Estimating training data influence by tracing gradient descent, 2020
Pruthi, G., Liu, F., Sundararajan, M., and Kale, S · 2002
Earlier work this paper cites.
Evaluation of similarity-based explanations, 2021
Hanawa, K., Yokoi, S., Hara, S., and Inui, K · 2006
Earlier work this paper cites.
Common voice: A massively-multilingual speech corpus
Ardila, R., Branson, M., Davis, K., Henretty, M., Kohler, M., Meyer, J., Morais, R., Saunders, L., Tyers, F. M., and Weber, G · 2020
Earlier work this paper cites.
Understanding black-box predictions via influence functions, 2020
Koh, P. W. and Liang, P · 2020
Earlier work this paper cites.
Articulating value from data
World Economic Forum · 2021
Earlier work this paper cites.
Training compute-optimal large language models, 2022
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., de Las Casas, D., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., van den Driessche, G., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Rae, J. W., Vinyals, O., and Sifre, L · 2022
Earlier work this paper cites.
Authors guild v. openai
Authors Guild · 2023
Earlier work this paper cites.
Ai is a lot of work
Dzieza, J · 2023
Earlier work this paper cites.
Datacomp: In search of the next generation of multimodal datasets, 2023
Gadre, S. Y., Ilharco, G., Fang, A., Hayase, J., Smyrnis, G., Nguyen, T., Marten, R., Wortsman, M., Ghosh, D., Zhang, J., Orgad, E., Entezari, R., Daras, G., Pratt, S., Ramanujan, V., Bitton, Y., Marathe, K., Mussmann, S., Vencu, R., Cherti, M., Krishna, R., Koh, P. W., Saukh, O., Ratner, A., Song, S., Hajishirzi, H., Farhadi, A., Beaumont, R., Oh, S., Dimakis, A., Jitsev, J., Carmon, Y., Shankar, V., and Schmidt, L · 2023
Earlier work this paper cites.
Openassistant conversations – democratizing large language model alignment, 2023
Köpf, A., Kilcher, Y., von Rütte, D., Anagnostidis, S., Tam, Z.-R., Stevens, K., Barhoum, A., Duc, N. M., Stanley, O., Nagyfi, R., ES, S., Suri, S., Glushkov, D., Dantuluri, A., Maguire, A., Schuhmann, C., Nguyen, H., and Mattick, A · 2023
Earlier work this paper cites.
Scaling data-constrained language models, 2023
Muennighoff, N., Rush, A. M., Barak, B., Scao, T. L., Piktus, A., Tazi, N., Pyysalo, S., Wolf, T., and Raffel, C · 2023
Earlier work this paper cites.
New york times v. openai
New York Times · 2023
Earlier work this paper cites.
Trak: Attributing model behavior at scale, 2023
Park, S. M., Georgiev, K., Ilyas, A., Leclerc, G., and Madry, A · 2023
Earlier work this paper cites.
Shutterstock expands partnership with openai, signs new six-year agreement to provide high-quality training data
Shutterstock · 2023
Cited alongside, same era.
Warstadt, A., Choshen, L., Mueller, A., Williams, A., Wilcox, E., and Zhuang, C · 2023
Cited alongside, same era.
Alden newspapers v. openai
Alden Newspapers · 2024
Cited alongside, same era.
What is your data worth to gpt? llm-scale data valuation with influence functions, 2024
Choe, S. K., Ahn, H., Bae, J., Zhao, K., Kang, M., Chung, Y., Pratapa, A., Neiswanger, W., Strubell, E., Mitamura, T., Schneider, J., Hovy, E., Grosse, R., and Xing, E · 2024
Cited alongside, same era.
Concord music group v. anthropic
Concord Music Group · 2024
Cited alongside, same era.
The fineweb datasets: Decanting the web for the finest text data at scale, 2024
Penedo, G., Kydlíček, H., allal, L. B., Lozhkov, A., Mitchell, M., Raffel, C., Werra, L. V., and Wolf, T · 2024
Later among the works it cites.
Reddit and openai build partnership
Reddit · 2024
Later among the works it cites.
Thomson reuters’ adjusted eps beats expectations, ai boosts results
Reuters · 2024
Later among the works it cites.
Dolma: an open corpus of three trillion tokens for language model pretraining research, 2024
Soldaini, L., Kinney, R., Bhagia, A., Schwenk, D., Atkinson, D., Authur, R., Bogin, B., Chandu, K., Dumas, J., Elazar, Y., Hofmann, V., Jha, A. H., Kumar, S., Lucy, L., Lyu, X., Lambert, N., Magnusson, I., Morrison, J., Muennighoff, N., Naik, A., Nam, C., Peters, M. E., Ravichander, A., Richardson, K., Shen, Z., Strubell, E., Subramani, N., Tafjord, O., Walsh, P., Zettlemoyer, L., Smith, N. A., Hajishirzi, H., Beltagy, I., Groeneveld, D., Dodge, J., and Lo, K · 2024
Later among the works it cites.
The atlantic announces product and content partnership with openai
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cottier, B., Rahman, R., Fattorini, L., Maslej, N., and Owen, D · 2024
Cited alongside, same era.
Language models scale reliably with over-training and on downstream tasks, 2024
Gadre, S. Y., Smyrnis, G., Shankar, V., Gururangan, S., Wortsman, M., Shao, R., Mercat, J., Fang, A., Li, J., Keh, S., Xin, R., Nezhurina, M., Vasiljevic, I., Jitsev, J., Soldaini, L., Dimakis, A. G., Ilharco, G., Koh, P. W., Song, S., Kollar, T., Carmon, Y., Dave, A., Heckel, R., Muennighoff, N., and Schmidt, L · 2024
Cited alongside, same era.
Releasing Common Corpus: the largest public domain dataset for training LLMs
Langlais, P.-C · 2024
Cited alongside, same era.
Datacomp-lm: In search of the next generation of training sets for language models
Li, J., Fang, A., Smyrnis, G., Ivgi, M., Jordan, M., Gadre, S., Bansal, H., Guha, E., Keh, S., Arora, K., et al · 2024
Cited alongside, same era.
Consent in crisis: The rapid decline of the ai data commons, 2024
Longpre, S., Mahari, R., Lee, A., Lund, C., Oderinwale, H., Brannon, W., Saxena, N., Obeng-Marnu, N., South, T., Hunter, C., Klyman, K., Klamm, C., Schoelkopf, H., Singh, N., Cherep, M., Anis, A., Dinh, A., Chitongo, C., Yin, D., Sileo, D., Mataciunas, D., Misra, D., Alghamdi, E., Shippole, E., Zhang, J., Materzynska, J., Qian, K., Tiwary, K., Miranda, L., Dey, M., Liang, M., Hamdy, M., Muennighoff, N., Ye, S., Kim, S., Mohanty, S., Gupta, V., Sharma, V., Chien, V. M., Zhou, X., Li, Y., Xiong, C., Villa, L., Biderman, S., Li, H., Ippolito, D., Hooker, S., Kabbara, J., and Pentland, S · 2024
Cited alongside, same era.
Silo language models: Isolating legal risk in a nonparametric datastore, 2024
Min, S., Gururangan, S., Wallace, E., Shi, W., Hajishirzi, H., Smith, N. A., and Zettlemoyer, L · 2024
Cited alongside, same era.
AI models that cost $1 billion to train are underway, $100 billion models coming — largest current models take “only” $100 million to train: Anthropic CEO
Morales, J · 2024
Cited alongside, same era.
The Atlantic · 2024
Later among the works it cites.
Will we run out of data? limits of llm scaling based on human-generated data, 2024
Villalobos, P., Ho, A., Sevilla, J., Besiroglu, T., Heim, L., and Hobbhahn, M · 2024
Later among the works it cites.
Vox media and openai form strategic content and product partnership
Vox Media · 2024
Later among the works it cites.
Breaking down emerging segments in the ai content licensing landscape
Wiley · 2024
Later among the works it cites.
Wildchat: 1m chatgpt interaction logs in the wild, 2024
Zhao, W., Ren, X., Hessel, J., Cardie, C., Choi, Y., and Deng, Y · 2024
Later among the works it cites.
https://huggingface.co/collections/r-three/common-pile-665a13e48528df6b00416dc0 , 2025
Common Pile · 2025
Closest in time.
Common crawl dataset
Common Crawl · 2025
Closest in time.
Data on notable ai models, 2024
Epoch AI · 2025
Closest in time.
Statistics on wages
International Labor Organization · 2025
Closest in time.
Html standard
W3C · 2025
Closest in time.