Fetching the paper…
Reading the bibliography…
Scaling the capacity of language models has consistently proven to be a reliable approach for improving performance and unlocking new capabilities.
Scaling laws for neural language models
J. Kaplan, S. McCandlish, T. Henighan, T. B. Brown, B. Chess, R. Child, S. Gray, A. Radford, J. Wu, and D. Amodei · 2001
Earlier work this paper cites.
Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
N. Shazeer, A. Mirhoseini, K. Maziarz, A. Davis, Q. Le, G. Hinton, and J. Dean · 2017
Earlier work this paper cites.
Know what you don’t know: Unanswerable questions for SQuAD
P. Rajpurkar, R. Jia, and P. Liang · 2018
Earlier work this paper cites.
Universal transformers
M. Dehghani, S. Gouws, O. Vinyals, J. Uszkoreit, and L. Kaiser · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova · 2019
Earlier work this paper cites.
CoQA: A conversational question answering challenge
S. Reddy, D. Chen, and C. D. Manning · 2019
Earlier work this paper cites.
Language models are few-shot learners
T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei · 2020
Earlier work this paper cites.
Scaling laws for autoregressive generative modeling
T. Henighan, J. Kaplan, M. Katz, M. Chen, C. Hesse, J. Jackson, H. Jun, T. B. Brown, P. Dhariwal, S. Gray, C. Hallacy, B. Mann, A. Radford, A. Ramesh, N. Ryder, D. M. Ziegler, J. Schulman, D. Amodei, and S. McCandlish · 2020
Earlier work this paper cites.
Training verifiers to solve math word problems
K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman · 2021
Earlier work this paper cites.
Glam: Efficient scaling of language models with mixture-of-experts
N. Du, Y. Huang, A. M. Dai, S. Tong, D. Lepikhin, Y. Xu, M. Krikun, Y. Zhou, A. W. Yu, O. Firat, B. Zoph, L. Fedus, M. Bosma, Z. Zhou, T. Wang, Y. E. Wang, K. Webster, M. Pellat, K. Robinson, K. S. Meier-Hellstern, T. Duke, L. Dixon, K. Zhang, Q. V. Le, Y. Wu, Z. Chen, and C. Cui · 2021
Earlier work this paper cites.
{GS}hard: Scaling giant models with conditional computation and automatic sharding
D. Lepikhin, H. Lee, Y. Xu, D. Chen, O. Firat, Y. Huang, M. Krikun, N. Shazeer, and Z. Chen · 2021
Earlier work this paper cites.
Gpt-neox-20b: An open-source autoregressive language model
S. Black, S. Biderman, E. Hallahan, Q. Anthony, L. Gao, L. Golding, H. He, C. Leahy, K. McDonell, J. Phang, et al · 2022
Earlier work this paper cites.
Unified scaling laws for routed language models
A. Clark, D. d. l. Casas, A. Guy, A. Mensch, M. Paganini, J. Hoffmann, B. Damoc, B. Hechtman, T. Cai, S. Borgeaud, G. v. d. Driessche, E. Rutherford, T. Hennigan, M. Johnson, K. Millican, A. Cassirer, C. Jones, E. Buchatskaya, D. Budden, L. Sifre, S. Osindero, O. Vinyals, J. Rae, E. Elsen, K. Kavukcuoglu, and K. Simonyan · 2022
Earlier work this paper cites.
Switch transformers: scaling to trillion parameter models with simple and efficient sparsity
W. Fedus, B. Zoph, and N. Shazeer · 2022
Earlier work this paper cites.
An empirical analysis of compute-optimal large language model training
J. Hoffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, T. Hennigan, E. Noland, K. Millican, G. van den Driessche, B. Damoc, A. Guy, S. Osindero, K. Simonyan, E. Elsen, O. Vinyals, J. Rae, and L. Sifre · 2022
Cited alongside, same era.
ST-MoE: designing stable and transferable sparse expert models
B. Zoph, I. Bello, S. Kumar, N. Du, Y. Huang, J. Dean, N. Shazeer, and W. Fedus · 2022
Cited alongside, same era.
J. Bai, S. Bai, Y. Chu, Z. Cui, K. Dang, X. Deng, Y. Fan, W. Ge, Y. Han, F. Huang, B. Hui, L. Ji, M. Li, J. Lin, R. Lin, D. Liu, G. Liu, C. Lu, K. Lu, J. Ma, R. Men, X. Ren, X. Ren, C. Tan, S. Tan, J. Tu, P. Wang, S. Wang, W. Wang, S. Wu, B. Xu, J. Xu, A. Yang, H. Yang, J. Yang, S. Yang, Y. Yao, B. Yu, H. Yuan, Z. Yuan, J. Zhang, X. Zhang, Y. Zhang, Z. Zhang, C. Zhou, J. Zhou, X. Zhou, and T. Zhu · 2023
Cited alongside, same era.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
BIG-bench authors · 2023
Cited alongside, same era.
Think before you speak: Training language models with pause tokens
S. Goyal, Z. Ji, A. S. Rawat, A. K. Menon, S. Kumar, and V. Nagarajan · 2024
Later among the works it cites.
X. O. He · 2024
Later among the works it cites.
Mixture of parrots: Experts improve memorization more than reasoning
S. Jelassi, C. Mohri, D. Brandfonbrener, A. Gu, N. Vyas, N. Anand, D. Alvarez-Melis, Y. Li, S. M. Kakade, and E. Malach · 2024
Later among the works it cites.
Q. Li, L. Cui, X. Zhao, L. Kong, and W. Bi · 2024
Later among the works it cites.
Scaling laws for fine-grained mixture of experts
J. Ludziejewski, J. Krajewski, K. Adamczewski, M. Pióro, M. Krutul, S. Antoniak, K. Ciebiera, K. Król, T. Odrzygóźdź, P. Sankowski, M. Cygan, and S. Jaszczur · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
MegaBlocks: Efficient Sparse Training with Mixture-of-Experts
T. Gale, D. Narayanan, C. Young, and M. Zaharia · 2023
Cited alongside, same era.
Gpt-4 technical report
OpenAI · 2023
Cited alongside, same era.
Redpajama: An open source recipe to reproduce llama training dataset
Together Computer · 2023
Cited alongside, same era.
Small-scale proxies for large-scale transformer training instabilities
M. Wortsman, P. J. Liu, L. Xiao, K. Everett, A. Alemi, B. Adlam, J. D. Co-Reyes, I. Gur, A. Kumar, R. Novak, et al · 2023
Cited alongside, same era.
Moeut: Mixture-of-experts universal transformers
R. Csord’as, K. Irie, J. Schmidhuber, C. Potts, and C. D. Manning · 2024
Cited alongside, same era.
Deepseek LLM: Scaling open-source language models with longtermism
DeepSeek-AI · 2024
Cited alongside, same era.
The llama 3 herd of models
A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. F. Amy Yan and, and et al · 2024
Cited alongside, same era.
Scaling laws for sparsely-connected foundation models
E. Frantar, C. R. Ruiz, N. Houlsby, D. Alistarh, and U. Evci · 2024
Cited alongside, same era.
Later among the works it cites.
Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models
I. Mirzadeh, K. Alizadeh, H. Shahrokhi, O. Tuzel, S. Bengio, and M. Farajtabar · 2024
Later among the works it cites.
Olmoe: Open mixture-of-experts language models, 2024
N. Muennighoff, L. Soldaini, D. Groeneveld, K. Lo, J. Morrison, S. Min, W. Shi, P. Walsh, O. Tafjord, N. Lambert, Y. Gu, S. Arora, A. Bhagia, D. Schwenk, D. Wadden, A. Wettig, B. Hui, T. Dettmers, D. Kiela, A. Farhadi, N. A. Smith, P. W. Koh, A. Singh, and H. Hajishirzi · 2024
Later among the works it cites.
Openai o1 system card
OpenAI · 2024
Later among the works it cites.
Scaling laws across model architectures: A comparative analysis of dense and MoE models in large language models
S. Wang, Z. Chen, B. Li, K. He, M. Zhang, and J. Wang · 2024
Later among the works it cites.
Toward inference-optimal mixture-of-expert large language models
L. Yun, Y. Zhuang, Y. Fu, E. P. Xing, and H. Zhang · 2024
Later among the works it cites.
A careful examination of large language model performance on grade school arithmetic
H. Zhang, J. Da, D. Lee, V. Robinson, C. Wu, W. Song, T. Zhao, P. Raja, D. Slack, Q. Lyu, et al · 2024
Later among the works it cites.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
DeepSeek-AI · 2025
Closest in time.
Nemo: a toolkit for conversational ai and large language models
NeMo Authors · 2025
Closest in time.