Fetching the paper…
Reading the bibliography…
We investigate the potential constraints on LLM scaling posed by the availability of public human-generated text data.
Exploring the limits of transfer learning with a unified text-to-text transformer
Raffel, C., Shazeer, N. M., Roberts, A., Lee, K., Narang, S., Matena, M., Zhou, Y., Li, W., and Liu, P. J · 1910
Earlier work this paper cites.
Optimal information processing and bayes’s theorem
Zellner, A · 1988
Earlier work this paper cites.
The size and growth rate of the internet, 1998
Coffman, K. and Odlyzko, A · 1998
Earlier work this paper cites.
Sizing the internet
Murray H., B. and Moore, A · 2000
Earlier work this paper cites.
The pushshift reddit dataset, 2020
Baumgartner, J., Zannettou, S., Keegan, B., Squire, M., and Blackburn, J · 2001
Earlier work this paper cites.
Scaling laws for neural language models, 2020
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2001
Earlier work this paper cites.
How much information, 2003
Lyman, P. and Varian, H. R · 2003
Earlier work this paper cites.
Language models are few-shot learners, 2020
Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I., and Amodei, D · 2005
Earlier work this paper cites.
An image is worth 16x16 words: Transformers for image recognition at scale, 2021
Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., Uszkoreit, J., and Houlsby, N · 2010
Earlier work this paper cites.
Scaling laws for autoregressive generative modeling, 2020
Henighan, T., Kaplan, J., Katz, M., Chen, M., Hesse, C., Jackson, J., Jun, H., Brown, T. B., Dhariwal, P., Gray, S., Hallacy, C., Mann, B., Radford, A., Ramesh, A., Ryder, N., Ziegler, D. M., Schulman, J., Amodei, D., and McCandlish, S · 2010
Earlier work this paper cites.
Average word length dynamics as indicator of cultural changes in society
Bochkarev, V. V., Shevlyakova, A. V., and Solovyev, V. D · 2012
Earlier work this paper cites.
Big data: Astronomical or genomical?
Stephens, Z. D., Lee, S., Faghri, F., Campbell, R. H., Zhai, C., Efron, M., Iyer, R. K., Schatz, M. C., Sinha, S., and Robinson, G. E · 2015
Earlier work this paper cites.
Investigation of the widely applicable bayesian information criterion, 2016
Friel, N., McKeone, J. P., Oates, C. J., and Pettitt, A. N · 2016
Earlier work this paper cites.
The growth rate and the nature of internet traffic
Odlyzko, A · 2016
Earlier work this paper cites.
Estimating search engine index size variability: A 9-year longitudinal study
van den Bosch, A., Bogers, T., and de Kunder, M · 2016
Earlier work this paper cites.
Cisco Visual Networking Index: Forecast and Methodology, 2016–2021, 2017
Cisco · 2017
Earlier work this paper cites.
Social emotion mining techniques for facebook posts reaction prediction, 2017
Krebs, F., Lubascher, B., Moers, T., Schaap, P., and Spanakis, G · 2017
Earlier work this paper cites.
Technology adoption
Ritchie, H. and Roser, M · 2017
Earlier work this paper cites.
Mastering chess and shogi by self-play with a general reinforcement learning algorithm, 2017
Silver, D., Hubert, T., Schrittwieser, J., Antonoglou, I., Lai, M., Guez, A., Lanctot, M., Sifre, L., Kumaran, D., Graepel, T., Lillicrap, T., Simonyan, K., and Hassabis, D · 2017
Earlier work this paper cites.
Ai safety via debate, 2018
Irving, G., Christiano, P., and Amodei, D · 2018
Earlier work this paper cites.
The digitization of the world from edge to core
Reinsel, D., Gantz, J., and Rydning, J · 2018
Earlier work this paper cites.
WhatsApp usage patterns and prediction of demographic characteristics without access to message content
Rosenfeld, A., Sina, S., Sarne, D., Avidov, O., and Kraus, S · 2018
Earlier work this paper cites.
Solving rubik’s cube with a robot hand, 2019
OpenAI, Akkaya, I., Andrychowicz, M., Chociej, M., Litwin, M., McGrew, B., Petron, A., Paino, A., Plappert, M., Powell, G., Ribas, R., Schneider, J., Tezak, N., Tworek, J., Welinder, P., Weng, L., Yuan, Q., Zaremba, W., and Zhang, L · 2019
Earlier work this paper cites.
Publications output: U.s. trends and international comparisons, 2019
White, K · 2019
Earlier work this paper cites.
The decay and persistence of web references
Loan, F. A. and Shah, U. Y · 2020
Earlier work this paper cites.
Number of people using the internet
Ritchie, H. and Roser, M · 2020
Earlier work this paper cites.
A large-scale corpus of E-mail conversations with standard and two-level dialogue act annotations
Taniguchi, M., Ueda, Y., Taniguchi, T., and Ohkuma, T · 2020
Earlier work this paper cites.
Exploring the limits of large scale pre-training, 2021
Abnar, S., Dehghani, M., Neyshabur, B., and Sedghi, H · 2021
Earlier work this paper cites.
Curriculum learning for language modeling
Campos, D · 2021
Earlier work this paper cites.
Measuring progress in deep reinforcement learning sample efficiency, 2021
Dorner, F. E · 2021
Earlier work this paper cites.
Glam: Efficient scaling of language models with mixture-of-experts, 2021
Du, N., Huang, Y., Dai, A. M., Tong, S., Lepikhin, D., Xu, Y., Krikun, M., Zhou, Y., Yu, A. W., Firat, O., Zoph, B., Fedus, L., Bosma, M., Zhou, Z., Wang, T., Wang, Y. E., Webster, K., Pellat, M., Robinson, K., Meier-Hellstern, K., Duke, T., Dixon, L., Zhang, K., Le, Q. V., Wu, Y., Chen, Z., and Cui, C · 2021
Earlier work this paper cites.
An empirical exploration in quality filtering of text data, 2021
Gao, L · 2021
Earlier work this paper cites.
The pile: An 800gb dataset of diverse text for language modeling, 2021
Gao, L., Biderman, S., Black, S., Golding, L., Hoppe, T., Foster, C., Phang, J., He, H., Thite, A., Nabeshima, N., Presser, S., and Leahy, C · 2021
Earlier work this paper cites.
Scaling laws for transfer, 2021
Hernandez, D., Kaplan, J., Henighan, T., and McCandlish, S · 2021
Earlier work this paper cites.
Gpt-3 powers the next generation of apps, 2021
OpenAI and Pilipiszyn, A · 2021
Earlier work this paper cites.
Mastering atari games with limited data, 2021
Ye, W., Liu, S., Kurutach, T., Abbeel, P., and Gao, Y · 2021
Earlier work this paper cites.
The paper of record meets an ephemeral web: An examination of linkrot and content drift within the new york times
Zittrain, J., Bowers, J., and Stanton, C · 2021
Earlier work this paper cites.
Constitutional AI: Harmlessness from AI Feedback, 2022
Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., Chen, A., Goldie, A., Mirhoseini, A., and McKinnon, C. et al · 2022
Earlier work this paper cites.
PaLM: Scaling Language Modeling with Pathways, 2022
Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., and Gehrmann, S. et al · 2022
Earlier work this paper cites.
Training compute-optimal large language models, 2022
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., Hennigan, T., Noland, E., Millican, K., Driessche, G. v. d., Damoc, B., Guy, A., Osindero, S., Simonyan, K., Elsen, E., Rae, J. W., Vinyals, O., and Sifre, L · 2022
Earlier work this paper cites.
The continued problem of url decay: an updated analysis of health care management journal citations
Howell, S. and Burtis, A. T · 2022
Earlier work this paper cites.
Large language models can self-improve, 2022
Huang, J., Gu, S. S., Hou, L., Wu, Y., Wang, X., Yu, H., and Han, J · 2022
Earlier work this paper cites.
Deduplicating training data makes language models better, 2022
Lee, K., Ippolito, D., Nystrom, A., Zhang, C., Eck, D., Callison-Burch, C., and Carlini, N · 2022
Earlier work this paper cites.
Quality not quantity: On the interaction between dataset design and robustness of clip, 2022
Nguyen, T., Ilharco, G., Wortsman, M., Oh, S., and Schmidt, L · 2022
Earlier work this paper cites.
chinchilla’s wild implications, 2022
Nostalgebraist · 2022
Earlier work this paper cites.
Gross domestic spending on R&D (indicator), 2022
OECD · 2022
Cited alongside, same era.
Reference hygiene and death on the internet–decay, rot, half-life, deterioration, and corruption
Ott, D. E · 2022
Cited alongside, same era.
Training language models to follow instructions with human feedback
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L. E., Simens, M., Askell, A., Welinder, P., Christiano, P. F., Leike, J., and Lowe, R. J · 2022
Cited alongside, same era.
Web citation analysis on journal of travel research: A study
Satyanarayana, D. and Damodar, P · 2022
Cited alongside, same era.
Self-critiquing models for assisting human evaluators, 2022
Saunders, W., Yeh, C., Wu, J., Bills, S., Ouyang, L., Ward, J., and Leike, J · 2022
Cited alongside, same era.
To repeat or not to repeat: Insights from scaling llm under token-crisis, 2023
Xue, F., Fu, Y., Zhou, W., Zheng, Z., and You, Y · 2023
Closest in time.
Leandojo: Theorem proving with retrieval-augmented language models, 2023
Yang, K., Swope, A. M., Gu, A., Chalamala, R., Song, P., Yu, S., Godil, S., Prenger, R., and Anandkumar, A · 2023
Closest in time.
Chain-of-thought reasoning is a policy improvement operator, 2023
Zhang, H. and Parkes, D. C · 2023
Closest in time.
A survey of large language models, 2023
Zhao, W. X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., Du, Y., Yang, C., Chen, Y., Chen, Z., Jiang, J., Ren, R., Li, Y., Tang, X., Liu, Z., Liu, P., Nie, J.-Y., and Wen, J.-R · 2023
Closest in time.
Ready-to-go transmission projects 2023
Zimmerman, Z., Goggin, M., and Gramlich, R · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Scao, T. L., Wang, T., Hesslow, D., Saulnier, L., Bekman, S., Bari, M. S., Biderman, S., Elsahar, H., Muennighoff, N., Phang, J., Press, O., Raffel, C., Sanh, V., Shen, S., Sutawika, L., Tae, J., Yong, Z. X., Launay, J., and Beltagy, I · 2022
Cited alongside, same era.
Compute trends across three eras of machine learning
Sevilla, J., Heim, L., Ho, A., Besiroglu, T., Hobbhahn, M., and Villalobos, P · 2022
Cited alongside, same era.
Galactica: A large language model for science, 2022
Taylor, R., Kardas, M., Cucurull, G., Scialom, T., Hartshorn, A., Saravia, E., Poulton, A., Kerkez, V., and Stojnic, R · 2022
Cited alongside, same era.
World population prospects 2022, online edition, 2022
United Nations · 2022
Cited alongside, same era.
Trends in training dataset sizes
Villalobos, P. and Ho, A · 2022
Cited alongside, same era.
Tensor programs v: Tuning large neural networks via zero-shot hyperparameter transfer
Yang, G., Hu, E. J., Babuschkin, I., Sidor, S., Liu, X., Farhi, D., Ryder, N., Pachocki, J., Chen, W., and Gao, J · 2022
Cited alongside, same era.
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., and Anadkat, S. et al · 2023
Cited alongside, same era.
Closest in time.
We knew the web was big…
Alpert, J. and Hajaj, N · 2024
Closest in time.
Projecting compute trends in machine learning, 2022
Besiroglu, T., Heim, L., and Sevilla, J · 2024
Closest in time.
Chinchilla scaling: A replication attempt, 2024
Besiroglu, T., Erdil, E., Barnett, M., and You, J · 2024
Closest in time.
Common crawl, 2024
Common Crawl · 2024
Closest in time.
Trends in the dollar training cost of machine learning systems, 2023
Cottier, B · 2024
Closest in time.
A tale of tails: Model collapse as a change of scaling laws, 2024
Dohmatob, E., Feng, Y., Yang, P., Charton, F., and Kempe, J · 2024
Closest in time.
Data Never Sleeps 10.0
Domo · 2024
Closest in time.
Parameter, compute and data trends in machine learning, 2022
Epoch · 2024
Closest in time.
Key trends and figures in machine learning, 2023
Epoch · 2024
Closest in time.
Visualizing Twitter’s Evolution 2012-2020 And How Tweeting Is Changing In The COVID-19 Era
GDELT · 2024
Closest in time.
Is model collapse inevitable? breaking the curse of recursion by accumulating real and synthetic data, 2024
Gerstgrasser, M., Schaeffer, R., Dey, A., Rafailov, R., Sleight, H., Hughes, J., Korbak, T., Agrawal, R., Pai, D., Gromov, A., Roberts, D. A., Yang, D., Donoho, D. L., and Koyejo, S · 2024
Closest in time.
Chatgpt creators openai are generating 100 billion words per day, ceo says, 2024
Griffin, A · 2024
Closest in time.
Algorithmic progress in language models
Ho, A., Besiroglu, T., Erdil, E., Owen, D., Rahman, R., Guo, Z. C., Atkinson, D., Thompson, N., and Sevilla, J · 2024
Closest in time.
Genomic data science (fact sheet), 2024
Institute, N. H. G. R · 2024
Closest in time.
GitHub - jerryspan/FacebookR: Facebook Post Reactions dataset
jerryspan · 2024
Closest in time.
Digital 2023: Global Overview Report — DataReportal
Kemp, S · 2024
Closest in time.
Facebook Messenger Users, Stats, Data, Trends, and More — DataReportal
Kemp, S · 2024
Closest in time.
Facebook Users, Stats, Data, Trends, and More — DataReportal
Kemp, S · 2024
Closest in time.
Debating with more persuasive llms leads to more truthful answers, 2024
Khan, A., Hughes, J., Valentine, D., Ruis, L., Sachan, K., Radhakrishnan, A., Grefenstette, E., Bowman, S. R., Rocktäschel, T., and Perez, E · 2024
Closest in time.
Large Text Compression Benchmark, 2006/2024
Mahoney, M · 2024
Closest in time.
Introducing Meta Llama 3: The most capable openly available LLM to date
Meta · 2024
Closest in time.
Introducing DBRX: A New State-of-the-Art Open LLM — Databricks — databricks.com
Mosaic AI · 2024
Closest in time.
Say :wave: to Messenger: Introducing New Messaging Features for Instagram
Mosseri, A. and Chudnovsky, S · 2024
Closest in time.
Indicators for the Presence of Languages in the Internet
OBDILCI · 2024
Closest in time.
Gpt-4v(ision) system card
OpenAI · 2024
Closest in time.
tiktoken, 2024
OpenAI · 2024
Closest in time.
Email statistics report, 2020-2024
Radicati · 2024
Closest in time.
Internet Live Stats
Real Time Statistics Project · 2024
Closest in time.
Google’s index size revealed: 400 billion docs (& changing), 2024
Shepard, C · 2024
Closest in time.
Nvidia to reportedly triple output of compute gpus in 2024: Up to 2 million h100s, Aug 2023
Shilov, A · 2024
Closest in time.
Top Websites
Similarweb · 2024
Closest in time.
WhatsApp is now delivering roughly 100 billion messages a day
Singh, M · 2024
Closest in time.
World wide web size, 2024
Size, W. W. W · 2024
Closest in time.
Solving olympiad geometry without human demonstrations
Trinh, T. H., Wu, Y., Le, Q. V., He, H., and Luong, T · 2024
Closest in time.
Wikipedia:Size of Wikipedia
Wikipedia · 2024
Closest in time.
Xverse-65b: A multilingual large language model, 2024
XVERSE Technology Inc · 2024
Closest in time.
Status of the CMOS Image Sensor Industry 2021
Yole Développement · 2024
Closest in time.
YouTube for Press
YouTube · 2024
Closest in time.
2021 worldwide image capture forecast: 2020 – 2025, 2021
Lee, E · 2025
Closest in time.