Fetching the paper…
Reading the bibliography…
Scaling laws for large language models (LLMs) predict model performance based on parameters like size and training data.
Hellaswag: Can a machine really finish your sentence?, 2019
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y · 1905
Earlier work this paper cites.
Winogrande: An adversarial winograd schema challenge at scale, 2019
Sakaguchi, K., Bras, R. L., Bhagavatula, C., and Choi, Y · 1907
Earlier work this paper cites.
A constructive prediction of the generalization error across scales, 2019
Rosenfeld, J. S., Rosenfeld, A., Belinkov, Y., and Shavit, N · 1909
Earlier work this paper cites.
Multivariate exploratory data analysis: A perspective on exploratory factor analysis
Yates, A · 1987
Earlier work this paper cites.
Monotonic networks
Sill, J · 1997
Earlier work this paper cites.
Scaling laws for neural language models, 2020
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D · 2001
Earlier work this paper cites.
Stochastic frontier analysis
Kumbhakar, S. C. and Lovell, C. K · 2003
Earlier work this paper cites.
Pattern recognition and machine learning , volume 4
Bishop, C. M. and Nasrabadi, N. M · 2006
Earlier work this paper cites.
18 multidimensional item response theory
Reckase, M. D · 2006
Earlier work this paper cites.
Multidimensional Item Response Theory
Reckase, M · 2009
Earlier work this paper cites.
Adam: A Method for Stochastic Optimization
Kingma, D. P. and Ba, J · 2017
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018
Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O · 2018
Earlier work this paper cites.
Glue: A multi-task benchmark and analysis platform for natural language understanding
Wang, A., Singh, A., Michael, J., Hill, F., Levy, O., and Bowman, S. R · 2018
Earlier work this paper cites.
Factor_analyzer documentation
Biggs, J · 2019
Earlier work this paper cites.
Joint maximum likelihood estimation for high-dimensional exploratory item factor analysis
Chen, Y., Li, X., and Zhang, S · 2019
Earlier work this paper cites.
Measuring Massive Multitask Language Understanding
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J · 2020
Cited alongside, same era.
Evaluating large language models trained on code
Chen, M., Tworek, J., Jun, H., Yuan, Q., de Oliveira Pinto, H. P., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., Ryder, N., Pavlov, M., Power, A., Kaiser, L., Bavarian, M., Winter, C., Tillet, P., Such, F. P., Cummings, D., Plappert, M., Chantzis, F., Barnes, E., Herbert-Voss, A., Guss, W. H., Nichol, A., Paino, A., Tezak, N., Tang, J., Babuschkin, I., Balaji, S., Jain, S., Saunders, W., Hesse, C., Carr, A. N., Leike, J., Achiam, J., Misra, V., Morikawa, E., Radford, A., Knight, M., Brundage, M., Murati, M., Mayer, K., Welinder, P., McGrew, B., Amodei, D., McCandlish, S., Sutskever, I., and Zaremba, W · 2021
Cited alongside, same era.
Training verifiers to solve math word problems, 2021
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., and Schulman, J · 2021
Cited alongside, same era.
Measuring mathematical problem solving with the math dataset, 2021
Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J · 2021
Musr: Testing the limits of chain-of-thought with multistep soft reasoning
Sprague, Z., Ye, X., Bostrom, K., Chaudhuri, S., and Durrett, G · 2023
Later among the works it cites.
Instruction-following evaluation for large language models, 2023
Zhou, J., Lu, T., Mishra, S., Brahma, S., Basu, S., Luan, Y., Zhou, D., and Hou, L · 2023
Later among the works it cites.
Physics of language models: Part 3.3, knowledge capacity scaling laws
Allen-Zhu, Z. and Li, Y · 2024
Closest in time.
Large language monkeys: Scaling inference compute with repeated sampling
Brown, B., Juravsky, J., Ehrlich, R., Clark, R., Le, Q. V., Ré, C., and Mirhoseini, A · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Truthfulqa: Measuring how models mimic human falsehoods
Lin, S., Hilton, J., and Evans, O · 2021
Cited alongside, same era.
Learning single-index models with shallow neural networks
Bietti, A., Bruna, J., Sanford, C., and Song, M. J · 2022
Cited alongside, same era.
Training compute-optimal large language models
Hoffmann, J., Borgeaud, S., Mensch, A., Buchatskaya, E., Cai, T., Rutherford, E., Casas, D. d. L., Hendricks, L. A., Welbl, J., Clark, A., et al · 2022
Cited alongside, same era.
Holistic evaluation of language models
Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., et al · 2022
Cited alongside, same era.
Challenging big-bench tasks and whether chain-of-thought can solve them, 2022
Suzgun, M., Scales, N., Schärli, N., Gehrmann, S., Tay, Y., Chung, H. W., Chowdhery, A., Le, Q. V., Chi, E. H., Zhou, D., and Wei, J · 2022
Cited alongside, same era.
Open llm leaderboard., 2023
Beeching, E., Fourrier, C., Habib, N., Han, S., Lambert, N., Rajani, N., Sanseviero, O., Tunstall, L., and Wolf, T · 2023
Cited alongside, same era.
Revealing the structure of language model capabilities
Burnell, R., Hao, H., Conway, A. R., and Orallo, J. H · 2023
Cited alongside, same era.
Unveiling the general intelligence factor in language models: A psychometric approach
Ilić, D · 2023
Cited alongside, same era.
Choshen, L., Zhang, Y., and Andreas, J · 2024
Closest in time.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Closest in time.
Open llm leaderboard v2
Fourrier, C., Habib, N., Lozovskaya, A., Szafer, K., and Wolf, T · 2024
Closest in time.
Language models scale reliably with over-training and on downstream tasks, 2024
Gadre, S. Y., Smyrnis, G., Shankar, V., Gururangan, S., Wortsman, M., Shao, R., Mercat, J., Fang, A., Li, J., Keh, S., Xin, R., Nezhurina, M., Vasiljevic, I., Jitsev, J., Soldaini, L., Dimakis, A. G., Ilharco, G., Koh, P. W., Song, S., Kollar, T., Carmon, Y., Dave, A., Heckel, R., Muennighoff, N., and Schmidt, L · 2024
Closest in time.
metabench - a sparse benchmark to measure general ability in large language models
Kipnis, A., Voudouris, K., Buschoff, L. M. S., and Schulz, E · 2024
Closest in time.
How predictable is language model benchmark performance?
Owen, D · 2024
Closest in time.
Eq-bench: An emotional intelligence benchmark for large language models, 2024
Paech, S. J · 2024
Closest in time.
Observational Scaling Laws and the Predictability of Language Model Performance, July 2024
Ruan, Y., Maddison, C. J., and Hashimoto, T · 2024
Closest in time.
Mmlu-pro: A more robust and challenging multi-task language understanding benchmark, 2024
Wang, Y., Ma, X., Zhang, G., Ni, Y., Chandra, A., Guo, S., Ren, W., Arulraj, A., He, X., Jiang, Z., Li, T., Ku, M., Wang, K., Zhuang, A., Fan, R., Yue, X., and Chen, W · 2024
Closest in time.
Yang, A., Yang, B., Hui, B., Zheng, B., Yu, B., Zhou, C., Li, C., Li, C., Liu, D., Huang, F., et al · 2024
Closest in time.
Yi: Open foundation models by 01. ai
Young, A., Chen, B., Li, C., Huang, C., Zhang, G., Zhang, G., Li, H., Zhu, J., Chen, J., Chang, J., et al · 2024
Closest in time.