Fetching the paper…
Reading the bibliography…
Kaplan et al.
A new method of interpolation and smooth curve fitting based on local procedures
H. Akima · 1970
Earlier work this paper cites.
J. L. Ba, J. R. Kiros, and G. E. Hinton · 2016
Earlier work this paper cites.
Accurate, large minibatch SGD: Training imagenet in 1 hour
P. Goyal, P. Dollár, R. Girshick, P. Noordhuis, L. Wesolowski, A. Kyrola, A. Tulloch, Y. Jia, and K. He · 2017
Earlier work this paper cites.
Deep learning scaling is predictable, empirically
J. Hestness, S. Narang, N. Ardalani, G. F. Diamos, H. Jun, H. Kianinejad, M. M. A. Patwary, Y. Yang, and Y. Zhou · 2017
Earlier work this paper cites.
Decoupled weight decay regularization
I. Loshchilov and F. Hutter · 2017
Earlier work this paper cites.
Attention is all you need
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin · 2017
Earlier work this paper cites.
An empirical model of large-batch training
S. McCandlish, J. Kaplan, D. Amodei, and The OpenAI Dota Team · 2018
Earlier work this paper cites.
Pytorch: An imperative style, high-performance deep learning library
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, et al · 2019
Earlier work this paper cites.
Language models are unsupervised multitask learners, 2019
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al · 2019
Earlier work this paper cites.
A constructive prediction of the generalization error across scales
J. S. Rosenfeld, A. Rosenfeld, Y. Belinkov, and N. Shavit · 2019
Earlier work this paper cites.
Measuring the effects of data parallelism on neural network training
C. J. Shallue, J. Lee, J. Antognini, J. Sohl-Dickstein, R. Frostig, and G. E. Dahl · 2019
Earlier work this paper cites.
Megatron-LM: Training multi-billion parameter language models using model parallelism
M. Shoeybi, M. Patwary, R. Puri, P. LeGresley, J. Casper, and B. Catanzaro · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al · 2020
Earlier work this paper cites.
The Pile: An 800gb dataset of diverse text for language modeling
L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, S. Presser, and C. Leahy · 2020
Earlier work this paper cites.
Scaling laws for autoregressive generative modeling
T. Henighan, J. Kaplan, M. Katz, M. Chen, C. Hesse, J. Jackson, H. Jun, T. B. Brown, P. Dhariwal, S. Gray, et al · 2020
Earlier work this paper cites.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2020
Earlier work this paper cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu · 2020
Earlier work this paper cites.
Glu variants improve transformer
N. Shazeer · 2020
Earlier work this paper cites.
Jurassic-1: Technical details and evaluation, 2021
O. Lieber, O. Sharir, B. Lenz, and Y. Shoham · 2021
Earlier work this paper cites.
Scaling language models: Methods, analysis & insights from training gopher
J. W. Rae, S. Borgeaud, T. Cai, K. Millican, J. Hoffmann, F. Song, J. Aslanides, S. Henderson, R. Ring, S. Young, et al · 2021
Cited alongside, same era.
Gpt-neox-20b: An open-source autoregressive language model
S. Black, S. Biderman, E. Hallahan, Q. Anthony, L. Gao, L. Golding, H. He, C. Leahy, K. McDonell, J. Phang, et al · 2022
Cited alongside, same era.
An empirical analysis of compute-optimal large language model training
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, et al · 2022
Cited alongside, same era.
Scaling laws under the microscope: Predicting transformer performance from small scale experiments
M. Ivgi, Y. Carmon, and J. Berant · 2022
Cited alongside, same era.
xformers: A modular and hackable transformer modelling library
B. Lefaudeux, F. Massa, D. Liskovich, W. Xiong, V. Caggiano, S. Naren, M. Xu, J. Hu, M. Tintore, S. Zhang, P. Labatut, D. Haziza, L. Wehrstedt, J. Reizenstein, and G. Sizov · 2022
URL https://www.lesswrong.com/posts/midXmMb2Xg37F2Kgn/new-scaling-laws-for-large-language-models
New scaling laws for large language models, 2023 · 2024
Closest in time.
Scaling mlps: A tale of inductive bias
G. Bachmann, S. Anagnostidis, and T. Hofmann · 2024
Closest in time.
Stable lm 2 1.6 b technical report
M. Bellagente, J. Tow, D. Mahan, D. Phung, M. Zhuravinskyi, R. Adithyan, J. Baicoianu, B. Brooks, N. Cooper, A. Datta, et al · 2024
Closest in time.
Chinchilla scaling: A replication attempt
T. Besiroglu, E. Erdil, M. Barnett, and J. You · 2024
Closest in time.
A dynamical model of neural scaling laws
B. Bordelon, A. Atanasov, and C. Pand Pehlevan · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A solvable model of neural scaling laws
A. Maloney, D. A. Roberts, and J. Sully · 2022
Cited alongside, same era.
BLOOM: A 176b-parameter open-access multilingual language model
T. L. Scao, A. Fan, C. Akiki, E. Pavlick, S. Ilić, D. Hesslow, R. Castagné, A. S. Luccioni, F. Yvon, et al · 2022
Cited alongside, same era.
Using DeepSpeed and Megatron to train Megatron-Turing NLG 530B, a large-scale generative language model
S. Smith, M. Patwary, B. Norick, P. LeGresley, S. Rajbhandari, J. Casper, Z. Liu, S. Prabhumoye, G. Zerveas, V. Korthikanti, et al · 2022
Cited alongside, same era.
Scale efficiently: Insights from pre-training and fine-tuning transformers
Y. Tay, M. Dehghani, J. Rao, W. Fedus, S. Abnar, H. W. Chung, S. Narang, D. Yogatama, A. Vaswani, and D. Metzler · 2022
Cited alongside, same era.
Scaling vision transformers
X. Zhai, A. Kolesnikov, N. Houlsby, and L. Beyer · 2022
Cited alongside, same era.
Getting ViT in shape: Scaling laws for compute-optimal model design
I. Alabdulmohsin, X. Zhai, A. Kolesnikov, and L. Beyer · 2023
Cited alongside, same era.
The onset of variance-limited behavior for networks in the lazy and rich regimes
A. Atanasov, B. Bordelon, S. Sainathan, and C. Pehlevan · 2023
Cited alongside, same era.
DeepSeek · 2024
Closest in time.
Language models scale reliably with over-training and on downstream tasks
S. Y. Gadre, G. Smyrnis, V. Shankar, S. Gururangan, M. Wortsman, R. Shao, J. Mercat, A. Fang, J. Li, S. Keh, R. Xin, M. Nezhurina, I. Vasiljevic, J. Jitsev, A. G. Dimakis, G. Ilharco, S. Song, T. Kollar, Y. Carmon, A. Dave, R. Heckel, N. Muennighoff, and L. Schmidt · 2024
Closest in time.
Scaling laws for data filtering–data curation cannot be compute agnostic
S. Goyal, P. Maini, Z. C. Lipton, A. Raghunathan, and J. Z. Kolter · 2024
Closest in time.
Scaling laws and compute-optimal training beyond fixed training durations
A. Hägele, E. Bakouch, A. Kosson, L. B. Allal, L. Von Werra, and M. Jaggi · 2024
Closest in time.
MiniCPM: Unveiling the potential of small language models with scalable training strategies
S. Hu, Y. Tu, X. Han, C. He, G. Cui, X. Long, Z. Zheng, Y. Fang, Y. Huang, W. Zhao, et al · 2024
Closest in time.
Information-theoretic foundations for neural scaling laws
H. J. Jeon and B. Van Roy · 2024
Closest in time.
Scaling laws in linear regression: Compute, parameters, and data
L. Lin, J. Wu, S. M. Kakade, P. L. Bartlett, and J. D. Lee · 2024
Closest in time.
Scaling data-constrained language models
N. Muennighoff, A. Rush, B. Barak, T. Le Scao, N. Tazi, A. Piktus, S. Pyysalo, T. Wolf, and C. A. Raffel · 2024
Closest in time.
4+3 phases of compute-optimal neural scaling laws
E. Paquette, C. Paquette, L. Xiao, and J. Pennington · 2024
Closest in time.
Reconciling Kaplan and Chinchilla scaling laws
T. Pearce and J. Song · 2024
Closest in time.
Are emergent abilities of large language models a mirage?
R. Schaeffer, B. Miranda, and S. Koyejo · 2024
Closest in time.
Roformer: Enhanced transformer with rotary position embedding
J. Su, M. Ahmed, Y. Lu, S. Pan, W. Bo, and Y. Liu · 2024
Closest in time.
Small-scale proxies for large-scale transformer training instabilities
M. Wortsman, P. J. Liu, L. Xiao, K. E. Everett, A. A. Alemi, B. Adlam, J. D. Co-Reyes, I. Gur, A. Kumar, R. Novak, J. Pennington, J. Sohl-Dickstein, K. Xu, J. Lee, J. Gilmer, and S. Kornblith · 2024
Closest in time.