Fetching the paper…
Reading the bibliography…
Past work has established scaling laws that predict the performance of a neural language model (LM) as a function of its parameter count and the number of tokens it's trained on, enabling optimal allocation of a fixed compute budget.
A mathematical theory of communication
Claude Elwood Shannon · 1948
Earlier work this paper cites.
Prediction and entropy of printed english
Claude E Shannon · 1951
Earlier work this paper cites.
A method for the construction of minimum-redundancy codes
David A Huffman · 1952
Earlier work this paper cites.
Toward the logical description of languages in their phonemic aspect
E Colin Cherry, Morris Halle, and Roman Jakobson · 1953
Earlier work this paper cites.
Three models for the description of language
Noam Chomsky · 1956
Earlier work this paper cites.
Syntax-controlled probabilities
Ulf Grenander · 1967
Earlier work this paper cites.
A universal algorithm for sequential data compression
Jacob Ziv and Abraham Lempel · 1977
Earlier work this paper cites.
A theory of language and information: a mathematical approach
Zellig Harris · 1991
Earlier work this paper cites.
Gnu gzip
Jean-loup Gailly and Mark Adler · 1992
Earlier work this paper cites.
Learning curves: Asymptotic values and rate of convergence
Corinna Cortes, Lawrence D Jackel, Sara Solla, Vladimir Vapnik, and John Denker · 1993
Earlier work this paper cites.
Deflate compressed data format specification version 1.3
Peter Deutsch · 1996
Earlier work this paper cites.
Statistical properties of probabilistic context-free grammars
Zhiyi Chi · 1999
Earlier work this paper cites.
NLTK: The natural language toolkit
Steven Bird and Edward Loper · 2004
Earlier work this paper cites.
Probabilistic context-free grammars estimated from infinite distributions
Anna Corazza and Giorgio Satta · 2007
Earlier work this paper cites.
Word lengths are optimized for efficient communication
Steven T Piantadosi, Harry Tily, and Edward Gibson · 2011
Earlier work this paper cites.
Sgdr: Stochastic gradient descent with warm restarts
Ilya Loshchilov and Frank Hutter · 2016
Cited alongside, same era.
Deep learning scaling is predictable, empirically
Joel Hestness, Sharan Narang, Newsha Ardalani, Gregory Diamos, Heewoo Jun, Hassan Kianinejad, Md Mostofa Ali Patwary, Yang Yang, and Yanqi Zhou · 2017
Cited alongside, same era.
Decoupled weight decay regularization
Ilya Loshchilov and Frank Hutter · 2017
Cited alongside, same era.
On the derivational entropy of left-to-right probabilistic finite-state automata and hidden Markov models
Joan Andreu Sánchez, Martha Alicia Rocha, Verónica Romero, and Mauricio Villegas · 2018
Cited alongside, same era.
Pytorch: An imperative style, high-performance deep learning library
Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al · 2019
Language modeling is compression
Grégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne, Elliot Catt, Tim Genewein, Christopher Mattern, Jordi Grau-Moya, Li Kevin Wenliang, Matthew Aitchison, Laurent Orseau, et al · 2023
Later among the works it cites.
Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, et al · 2023
Later among the works it cites.
“low-resource” text classification: A parameter-free classification method with compressors
Zhiying Jiang, Matthew Yang, Mikhail Tsirlin, Raphael Tang, Yiqin Dai, and Jimmy Lin · 2023
Later among the works it cites.
Starcoder: may the source be with you!
Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, et al · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A constructive prediction of the generalization error across scales
Jonathan S Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, and Nir Shavit · 2019
Cited alongside, same era.
Scaling laws for neural language models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei · 2020
Cited alongside, same era.
thomasbreydo/pcfg
Thomas Breydo · 2021
Cited alongside, same era.
Scaling scaling laws with board games
Andy L Jones · 2021
Cited alongside, same era.
Scaling language models: Methods, analysis & insights from training gopher
Jack W Rae, Sebastian Borgeaud, Trevor Cai, Katie Millican, Jordan Hoffmann, Francis Song, John Aslanides, Sarah Henderson, Roman Ring, Susannah Young, et al · 2021
Cited alongside, same era.
Estimating the entropy of linguistic distributions
Aryaman Arora, Clara Meister, and Ryan Cotterell · 2022
Cited alongside, same era.
Training compute-optimal large language models
Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, et al · 2022
Cited alongside, same era.
Ziming Liu and Max Tegmark · 2023
Later among the works it cites.
Scaling data-constrained language models
Niklas Muennighoff, Alexander Rush, Boaz Barak, Teven Le Scao, Nouamane Tazi, Aleksandra Piktus, Sampo Pyysalo, Thomas Wolf, and Colin A Raffel · 2023
Later among the works it cites.
Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay · 2023
Later among the works it cites.
Pytorch fsdp: experiences on scaling fully sharded data parallel
Yanli Zhao, Andrew Gu, Rohan Varma, Liang Luo, Chien-Chin Huang, Min Xu, Less Wright, Hamid Shojanazeri, Myle Ott, Sam Shleifer, et al · 2023
Later among the works it cites.
Deepseek llm: Scaling open-source language models with longtermism
Xiao Bi, Deli Chen, Guanting Chen, Shanhuang Chen, Damai Dai, Chengqi Deng, Honghui Ding, Kai Dong, Qiushi Du, Zhe Fu, et al · 2024
Closest in time.
Compression represents intelligence linearly
Yuzhen Huang, Jinghan Zhang, Zifei Shan, and Junxian He · 2024
Closest in time.
Mission: Impossible language models
Julie Kallini, Isabel Papadimitriou, Richard Futrell, Kyle Mahowald, and Christopher Potts · 2024
Closest in time.
Fineweb, 2024
Guilherme Penedo, Hynek Kydlíček, Leandro von Werra, and Thomas Wolf · 2024
Closest in time.
Doremi: Optimizing data mixtures speeds up language model pretraining
Sang Michael Xie, Hieu Pham, Xuanyi Dong, Nan Du, Hanxiao Liu, Yifeng Lu, Percy S Liang, Quoc V Le, Tengyu Ma, and Adams Wei Yu · 2024
Closest in time.
Are all languages equally hard to language-model?
Ryan Cotterell, Sabrina J. Mielke, Jason Eisner, and Brian Roark · 2085
Closest in time.