Fetching the paper…
Reading the bibliography…
Hoffmann et al.
The large-sample distribution of the likelihood ratio for testing composite hypotheses
Wilks, S. S. (1938) · 1938
Earlier work this paper cites.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al. (2022) · 2022
Earlier work this paper cites.
Anil, R., Dai, A. M., Firat, O., Johnson, M., Lepikhin, D., Passos, A., Shakeri, S., Taropa, E., Bailey, P., Chen, Z., et al. (2023) · 2023
Earlier work this paper cites.
Gemini: a family of highly capable multimodal models
Google, G. T. (2023) · 2023
Cited alongside, same era.
Beyond chinchilla-optimal: Accounting for inference in language model scaling laws
Sardana, N. and Frankle, J. (2023) · 2023
Cited alongside, same era.
Deepseek llm: Scaling open-source language models with longtermism
Bi, X., Chen, D., Chen, G., Chen, S., Dai, D., Deng, C., Ding, H., Dong, K., Du, Q., Fu, Z., et al. (2024) · 2024
Cited alongside, same era.
A dynamical model of neural scaling laws
Bordelon, B., Atanasov, A., and Pehlevan, C. (2024) · 2024
Closest in time.
The quantization model of neural scaling
Michaud, E., Liu, Z., Girit, U., and Tegmark, M. (2024) · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…