Fetching the paper…
Reading the bibliography…
Ensuring safe and effective use of AI requires understanding and anticipating its performance on novel tasks, from advanced scientific challenges to transformed workplace activities.
Language models are few-shot learners
Brown, T., Mann, B., Ryder, N., Subbiah, M., Kaplan, J. D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al. (2020) · 1901
Earlier work this paper cites.
Hellaswag: Can a machine really finish your sentence?
Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. (2019) · 1905
Earlier work this paper cites.
Language tasks and language games: On methodology in current natural language processing research
Schlangen, D. (2019) · 1908
Earlier work this paper cites.
An outline of individual study
Partridge, G. E. (1910) · 1910
Earlier work this paper cites.
Ability, motivation, and speed
Thurstone, L. L. (1937) · 1937
Earlier work this paper cites.
Primary mental abilities: Psychometric monographs no. 1
Thurstone, L. L. (1938) · 1938
Earlier work this paper cites.
Marks of readable style; a study in adult education
Flesch, R. (1943) · 1943
Earlier work this paper cites.
On the theory of scales of measurement
Stevens, S. S. (1946) · 1946
Earlier work this paper cites.
The linear logistic test model as an instrument in educational research
Fischer, G. H. (1973) · 1973
Earlier work this paper cites.
The delphi method
Linstone, H. A., Turoff, M., et al. (1975) · 1975
Earlier work this paper cites.
The ‘ability’scale in item characteristic curve theory
Lord, F. M. (1975) · 1975
Earlier work this paper cites.
A study of computer-administered stradaptive ability testing
Vale, C. D. and Weiss, D. J. (1975) · 1975
Earlier work this paper cites.
Person reliability
Lumsden, J. (1977) · 1977
Earlier work this paper cites.
Development of a scale of cognitive demand for analysis of printed secondary science materials
Edwards, J. and Dall’Alba, G. (1981) · 1981
Earlier work this paper cites.
The person response curve: Fit of individuals to item response theory models
Trabin, T. E. and Weiss, D. J. (1983) · 1983
Earlier work this paper cites.
Estimating within-group interrater reliability with and without response bias
James, L. R., Demaree, R. G., and Wolf, G. (1984) · 1984
Earlier work this paper cites.
Unidimensional irt calibration of compensatory and noncompensatory multidimensional items
Ackerman, T. A. (1989) · 1989
Earlier work this paper cites.
The development of a tool for gauging the demands of gcse and a level exam questions
Hughes, S., Pollitt, A., and Ahmed, A. (1998) · 1998
Earlier work this paper cites.
Measurement in psychology: A critical history of a methodological concept
Michell, J. (1999) · 1999
Earlier work this paper cites.
Item response theory for psychologists
Embretson, S. and Reise, S. (2000) · 2000
Earlier work this paper cites.
The basics of item response theory
Baker, F. B. (2001) · 2001
Earlier work this paper cites.
Random forests
Breiman, L. (2001) · 2001
Earlier work this paper cites.
A revision of Bloom’s taxonomy: An overview
Krathwohl, D. R. (2002) · 2002
Earlier work this paper cites.
Test scoring
Thissen, D. and Wainer, H. (2002) · 2002
Earlier work this paper cites.
Explanatory item response models: A generalized linear and nonlinear approach
De Boeck, P. (2004) · 2004
Earlier work this paper cites.
A class of models for cognitive diagnosis
von Davier, M. and Yamamoto, K. (2004) · 2004
Earlier work this paper cites.
The cattell-horn-carroll theory of cognitive abilities: Past, present, and future
McGrew, K. S. (2005) · 2005
Earlier work this paper cites.
Chapter 18 - multidimensional item response theory
Reckase, M. D. (2006) · 2006
Earlier work this paper cites.
Comparative cognition: Experimental explorations of animal intelligence
Wasserman, E. A. and Zentall, T. R. (2006) · 2006
Earlier work this paper cites.
The demands of examination syllabuses and question papers
Pollitt, A., Ahmed, A., and Crisp, V. (2007) · 2007
Earlier work this paper cites.
Answers to 20 questions about interrater reliability and interrater agreement
LeBreton, J. M. and Senter, J. L. (2008) · 2008
Earlier work this paper cites.
A general diagnostic model applied to language testing data
Von Davier, M. (2008) · 2008
Earlier work this paper cites.
Is this year’s exam as demanding as last year’s? using a pilot method to evaluate the consistency of examination demands over time
Crisp, V. and Novaković, N. z. d. (2009) · 2009
Earlier work this paper cites.
Theory and practice of item response theory
De Ayala, R. J. (2009) · 2009
Earlier work this paper cites.
Coh-metrix: Providing multilevel analyses of text characteristics
Graesser, A. C., McNamara, D. S., and Kulikowich, J. M. (2011) · 2011
Earlier work this paper cites.
Cognitive load theory: How many types of load does it really need?
Kalyuga, S. (2011) · 2011
Earlier work this paper cites.
Scikit-learn: Machine learning in python
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., et al. (2011) · 2011
Earlier work this paper cites.
Cognitive load theory
Sweller, J. (2011) · 2011
Earlier work this paper cites.
Mapping the landscape of human-level artificial general intelligence
Adams, S., Arel, I., Bach, J., Coop, R., Furlan, R., Goertzel, B., Hall, J. S., Samsonovich, A., Scheutz, M., Schlesinger, M., et al. (2012) · 2012
Earlier work this paper cites.
Do nlp and machine learning improve traditional readability formulas?
François, T. and Miltsakaki, E. (2012) · 2012
Earlier work this paper cites.
When can categorical variables be treated as continuous? a comparison of robust continuous and categorical sem estimation methods under suboptimal conditions
Rhemtulla, M., Brosseau-Liard, P. É., and Savalei, V. (2012) · 2012
Earlier work this paper cites.
Item response theory
Embretson, S. E. and Reise, S. P. (2013) · 2013
Earlier work this paper cites.
Test validity
Wainer, H. and Braun, H. I. (2013) · 2013
Earlier work this paper cites.
Multidimensional explanatory item response modeling
De Boeck, P. and Wilson, M. (2014) · 2014
Earlier work this paper cites.
A general approach for assessing person fit and person reliability in typical-response measurement
Ferrando, P. J. (2014) · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Pennington, J., Socher, R., and Manning, C. D. (2014) · 2014
Earlier work this paper cites.
Using hierarchical irt models to create unidimensional measures from multidimensional data
Stucky, B. D. and Edelen, M. O. (2014) · 2014
Earlier work this paper cites.
A family of generalized diagnostic classification models for multiple choice option-based scoring
DiBello, L. V., Henson, R. A., and Stout, W. F. (2015) · 2015
Earlier work this paper cites.
Measurement: A very short introduction
Hand, D. J. (2016) · 2016
Earlier work this paper cites.
Building an evaluation scale using item response theory
Lalor, J. P., Wu, H., and Yu, H. (2016) · 2016
Earlier work this paper cites.
Handbook of item response theory
Van der Linden, W. J. and van der Linden, W. (2016) · 2016
Earlier work this paper cites.
Bifactor models in psychometric test development
Chen, F. F. and Zhang, Z. (2018) · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Radford, A., Narasimhan, K., Salimans, T., and Sutskever, I. (2018) · 2018
Earlier work this paper cites.
Multidimensional item response theory
Bonifay, W. (2019) · 2019
Earlier work this paper cites.
Rasch and rationality: Scale typologies as applied to item response theory
Freund, R. (2019) · 2019
Earlier work this paper cites.
We need to talk about standard splits
Gorman, K. and Bedrick, S. (2019) · 2019
Earlier work this paper cites.
AI extenders: the ethical and societal implications of humans cognitively extended by AI
Hernández-Orallo, J. and Vold, K. (2019) · 2019
Earlier work this paper cites.
Item response theory in ai: Analysing machine learning classifiers at the instance level
Martínez-Plumed, F., Prudêncio, R. B., Martínez-Usó, A., and Hernández-Orallo, J. (2019) · 2019
Earlier work this paper cites.
Machine behaviour
Rahwan, I., Cebrian, M., Obradovich, N., Bongard, J., Bonnefon, J.-F., Breazeal, C., Crandall, J. W., Christakis, N. A., Couzin, I. D., Jackson, M. O., et al. (2019) · 2019
Earlier work this paper cites.
Introduction to folk psychology: Pluralistic approaches
Andrews, K., Spaulding, S., and Westra, E. (2020) · 2020
Earlier work this paper cites.
Ai evaluation: On broken yardsticks and measurement scales
Hernandez-Orallo, J. (2020) · 2020
Cited alongside, same era.
Sim2real predictivity: Does evaluation in simulation predict real-world performance?
Kadian, A., Truong, J., Gokaslan, A., Clegg, A., Wijmans, E., Lee, S., Savva, M., Chernova, S., and Batra, D. (2020) · 2020
Cited alongside, same era.
Transformers: State-of-the-art natural language processing
Wolf, T., Debut, L., Sanh, V., Chaumond, J., Delangue, C., Moi, A., Cistac, P., Rault, T., Louf, R., Funtowicz, M., et al. (2020) · 2020
Cited alongside, same era.
Item response theory
Bock, R. D. and Gibbons, R. D. (2021) · 2021
Cited alongside, same era.
What will it take to fix benchmarking in natural language understanding?
Bowman, S. R. and Dahl, G. E. (2021) · 2021
Cited alongside, same era.
EU Artificial Intelligence Act
Automated evaluation of retrieval-augmented language models with task-specific exam generation
Guinet, G., Omidvar-Tehrani, B., Deoras, A., and Callot, L. (2024) · 2024
Later among the works it cites.
More than marketing? on the information value of ai benchmarks for practitioners
Hardy, A., Reuel, A., Meimandi, K. J., Soder, L., Griffith, A., Asmar, D. M., Koyejo, S., Bernstein, M. S., and Kochenderfer, M. J. (2024) · 2024
Later among the works it cites.
Machine learning with a reject option: A survey
Hendrickx, K., Perini, L., Van der Plas, D., Meert, W., and Davis, J. (2024) · 2024
Later among the works it cites.
Caveats and solutions for characterising general-purpose AI
Hernandez-Orallo, J. (2024) · 2024
Later among the works it cites.
Auxiliary task demands mask the capabilities of smaller language models
Hu, J. and Frank, M. (2024) · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
European Union (2024) · 2021
Cited alongside, same era.
General intelligence disentangled via a generality metric for natural and artificial intelligence
Hernández-Orallo, J., Loe, B. S., Cheke, L., Martínez-Plumed, F., and Ó hÉigeartaigh, S. (2021) · 2021
Cited alongside, same era.
Aleatoric and epistemic uncertainty in machine learning: An introduction to concepts and methods
Hüllermeier, E. and Waegeman, W. (2021) · 2021
Cited alongside, same era.
Dynabench: Rethinking benchmarking in nlp
Kiela, D., Bartolo, M., Nie, Y., Kaushik, D., Geiger, A., Wu, Z., Vidgen, B., Prasad, G., Singh, A., Ringshia, P., et al. (2021) · 2021
Cited alongside, same era.
Are we learning yet? a meta review of evaluation failures across machine learning
Liao, T., Taori, R., Raji, D., and Schmidt, L. (2021) · 2021
Cited alongside, same era.
Modern psychometrics: The science of psychological assessment
Rust, J., Kosinski, M., and Stillwell, D. (2021) · 2021
Cited alongside, same era.
Predicting the performance of multilingual nlp models
Srinivasan, A., Sitaram, S., Ganu, T., Dandapat, S., Bali, K., and Choudhury, M. (2021) · 2021
Cited alongside, same era.
Hurst, A., Lerer, A., Goucher, A. P., Perelman, A., Ramesh, A., Clark, A., Ostrow, A., Welihinda, A., Hayes, A., Radford, A., et al. (2024) · 2024
Later among the works it cites.
Evidence of interrelated cognitive-like capabilities in large language models: Indications of artificial general intelligence or achievement?
Ilić, D. and Gignac, G. E. (2024) · 2024
Later among the works it cites.
Jaech, A., Kalai, A., Lerer, A., Richardson, A., El-Kishky, A., Low, A., Helyar, A., Madry, A., Beutel, A., Carney, A., et al. (2024) · 2024
Later among the works it cites.
Does data contamination make a difference? insights from intentionally contaminating pre-training data for language models
Jiang, M., Liu, K., Zhong, M., Schaeffer, R., Ouyang, S., Han, J., and Koyejo, S. (2024a) · 2024
Later among the works it cites.
Medcalc-bench: Evaluating large language models for medical calculations
Khandekar, N., Jin, Q., Xiong, G., Dunn, S., Applebaum, S. S., Anwar, Z., Sarfo-Gyamfi, M., Safranek, C. W., Anwar, A. A., Zhang, A., et al. (2024) · 2024
Later among the works it cites.
Kipnis, A., Voudouris, K., Buschoff, L. M. S., and Schulz, E. (2024) · 2024
Later among the works it cites.
Item response theory for natural language processing
Lalor, J. P., Rodriguez, P., Sedoc, J., and Hernandez-Orallo, J. (2024) · 2024
Later among the works it cites.
Evaluating human-language model interaction
Lee, M., Srivastava, M., Hardy, A., Thickstun, J., Durmus, E., Paranjape, A., Gerard-Ursin, I., Li, X. L., Ladhak, F., Rong, F., et al. (2024) · 2024
Later among the works it cites.
GAOKAO-eval: Does high scores truly reflect strong capabilities in LLMs?
Lei, Z., Liang, T., Hu, H., Zhang, J., Zhou, Y., Shao, Y., Li, L., Li, C., Wang, C., Yan, H., and Guo, Q. (2024) · 2024
Later among the works it cites.
Same task, more tokens: the impact of input length on the reasoning performance of large language models
Levy, M., Jacoby, A., and Goldberg, Y. (2024) · 2024
Later among the works it cites.
ECBD: Evidence-centered benchmark design for NLP
Liu, Y. L., Blodgett, S. L., Cheung, J., Liao, Q. V., Olteanu, A., and Xiao, Z. (2024) · 2024
Later among the works it cites.
Inadequacies of large language model benchmarks in the era of generative artificial intelligence
McIntosh, T. R., Susnjak, T., Liu, T., Watters, P., and Halgamuge, M. N. (2024) · 2024
Later among the works it cites.
Gsm-symbolic: Understanding the limitations of mathematical reasoning in large language models
Mirzadeh, I., Alizadeh, K., Shahrokhi, H., Tuzel, O., Bengio, S., and Farajtabar, M. (2024) · 2024
Later among the works it cites.
Liar, liar, logical mire: A benchmark for suppositional reasoning in large language models
Mondorf, P. and Plank, B. (2024) · 2024
Later among the works it cites.
Language task difficulty prediction through LLM-annotated meta-features
Moros-Daval, Y., Martínez-Plumed, F., and Hernández-Orallo, J. (2024) · 2024
Later among the works it cites.
Education at a glance 2024: Oecd indicators
OECD (2024) · 2024
Later among the works it cites.
Routellm: Learning to route llms from preference data
Ong, I., Almahairi, A., Wu, V., Chiang, W.-L., Wu, T., Gonzalez, J. E., Kadous, M. W., and Stoica, I. (2024) · 2024
Later among the works it cites.
100 instances is all you need: predicting the success of a new LLM on unseen data by testing on a few instances
Pacchiardi, L., Cheke, L. G., and Hernández-Orallo, J. (2024) · 2024
Later among the works it cites.
Evaluating frontier models for dangerous capabilities
Phuong, M., Aitchison, M., Catt, E., Cogan, S., Kaskasoli, A., Krakovna, V., Lindner, D., Rahtz, M., Assael, Y., Hodkinson, S., et al. (2024) · 2024
Later among the works it cites.
tinybenchmarks: evaluating LLMs with fewer examples
Polo, F. M., Weber, L., Choshen, L., Sun, Y., Xu, G., and Yurochkin, M. (2024) · 2024
Later among the works it cites.
Gaps in the safety evaluation of generative ai
Rauh, M., Marchal, N., Manzini, A., Hendricks, L. A., Comanescu, R., Akbulut, C., Stepleton, T., Mateos-Garcia, J., Bergman, S., Kay, J., et al. (2024) · 2024
Later among the works it cites.
Safetywashing: Do AI safety benchmarks actually measure safety progress?
Ren, R., Basart, S., Khoja, A., Gatti, A., Phan, L., Yin, X., Mazeika, M., Pan, A., Mukobi, G., Kim, R. H., Fitz, S., and Hendrycks, D. (2024) · 2024
Later among the works it cites.
BetterBench: Assessing AI benchmarks, uncovering issues, and establishing best practices
Reuel, A., Hardy, A., Smith, C., Lamparth, M., Hardy, M., and Kochenderfer, M. (2024) · 2024
Later among the works it cites.
Observational scaling laws and the predictability of language model performance
Ruan, Y., Maddison, C. J., and Hashimoto, T. (2024) · 2024
Later among the works it cites.
Shorinwa, O., Mei, Z., Lidard, J., Ren, A. Z., and Majumdar, A. (2024) · 2024
Later among the works it cites.
Examining the robustness of LLM evaluation to the distributional assumptions of benchmarks
Siska, C., Marazopoulou, K., Ailem, M., and Bono, J. (2024) · 2024
Later among the works it cites.
Functional benchmarks for robust evaluation of reasoning performance, and the reasoning gap
Srivastava, S., PV, A., Menon, S., Sukumar, A., Philipose, A., Prince, S., Thomas, S., et al. (2024) · 2024
Later among the works it cites.
Do large language models perform the way people expect? measuring the human generalization function
Vafa, K., Rambachan, A., and Mullainathan, S. (2024) · 2024
Later among the works it cites.
Ai sandbagging: Language models can strategically underperform on evaluations
van der Weij, T., Hofstätter, F., Jaffe, O., Brown, S. F., and Ward, F. R. (2024) · 2024
Later among the works it cites.
Introducing flexible monotone multiple choice item response theory models and bit scales
Wallmark, J., Josefsson, M., and Wiberg, M. (2024) · 2024
Later among the works it cites.
Wang, Y. and Zhao, Y. (2024) · 2024
Later among the works it cites.
Livebench: A challenging, contamination-free llm benchmark
White, C., Dooley, S., Roberts, M., Pal, A., Feuer, B., Jain, S., Shwartz-Ziv, R., Jain, N., Saifullah, K., Naidu, S., et al. (2024) · 2024
Later among the works it cites.
Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., et al. (2024) · 2024
Later among the works it cites.
Zeng, Y. (2024) · 2024
Later among the works it cites.
Zhang, J., Huang, W., Ma, Z., Michel, O., He, D., Gupta, T., Ma, W.-C., Farhadi, A., Kembhavi, A., and Krishna, R. (2024) · 2024
Later among the works it cites.
An llm feature-based framework for dialogue constructiveness assessment
Zhou, L., Farag, Y., and Vlachos, A. (2024a) · 2024
Later among the works it cites.
Dynamic evaluation of large language models by meta probing agents
Zhu, K., Wang, J., Zhao, Q., Xu, R., and Xie, X. (2024) · 2024
Later among the works it cites.
From static benchmarks to adaptive testing: Psychometrics in AI evaluation
Zhuang, Y., Liu, Q., Ning, Y., Huang, W., Pardos, Z. A., Kyllonen, P. C., Zu, J., Mao, Q., Lv, R., Huang, Z., Zhao, G., Zhang, Z., Wang, S., and Chen, E. (2024) · 2024
Later among the works it cites.
Balepur, N., Rudinger, R., and Boyd-Graber, J. L. (2025) · 2025
Closest in time.
Open problems in machine unlearning for ai safety
Barez, F., Fu, T., Prabhu, A., Casper, S., Sanyal, A., Bibi, A., O’Gara, A., Kirk, R., Bucknall, B., Fist, T., et al. (2025) · 2025
Closest in time.
Paradigms of AI evaluation: Mapping goals, methodologies and culture
Burden, J., Tešić, M., Pacchiardi, L., and Hernández-Orallo, J. (2025) · 2025
Closest in time.
Can we trust ai benchmarks? an interdisciplinary review of current issues in ai evaluation
Eriksson, M., Purificato, E., Noroozian, A., Vinagre, J., Chaslot, G., Gomez, E., and Fernandez-Llorca, D. (2025) · 2025
Closest in time.
Improving LLM leaderboards with psychometrical methodology
Federiakin, D. (2025) · 2025
Closest in time.
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning
Guo, D., Yang, D., Zhang, H., Song, J., Zhang, R., Xu, R., Zhu, Q., Ma, S., Wang, P., Bi, X., et al. (2025) · 2025
Closest in time.
Kazemi, M., Fatemi, B., Bansal, H., Palowitch, J., Anastasiou, C., Mehta, S. V., Jain, L. K., Aglietti, V., Jindal, D., Chen, P., et al. (2025) · 2025
Closest in time.
The ethical evaluation of large language models and its optimization
Lyu, Y. and Du, Y. (2025) · 2025
Closest in time.
Alignvlm: Bridging vision and language latent spaces for multimodal understanding
Masry, A., Rodriguez, J. A., Zhang, T., Wang, S., Wang, C., Feizi, A., Suresh, A. K., Puri, A., Jian, X., Noël, P.-A., et al. (2025) · 2025
Closest in time.
PredictaBoard: Benchmarking LLM score predictability
Pacchiardi, L., Voudouris, K., Slater, B., Martínez-Plumed, F., Hernández-Orallo, J., Zhou, L., and Schellaert, W. (2025) · 2025
Closest in time.
Phan, L., Gatti, A., Han, Z., Li, N., Hu, J., Zhang, H., Shi, S., Choi, M., Agrawal, A., Chopra, A., et al. (2025) · 2025
Closest in time.
The Evaluation of Artificial Intelligence as a Prediction Problem
Schellaert, W. (2025) · 2025
Closest in time.
Analysing the predictability of language model performance
Schellaert, W., Martínez-Plumed, F., and Hernández-Orallo, J. (2025) · 2025
Closest in time.
Visual cognition in multimodal large language models
Schulze Buschoff, L. M., Akata, E., Bethge, M., and Schulz, E. (2025) · 2025
Closest in time.
What large language models know and what people think they know
Steyvers, M., Tejeda, H., Kumar, A., Belem, C., Karny, S., Hu, X., Mayer, L. W., and Smyth, P. (2025) · 2025
Closest in time.
Value compass leaderboard: A platform for fundamental and validated evaluation of llms values
Yao, J., Yi, X., Duan, S., Wang, J., Bai, Y., Huang, M., Zhang, P., Lu, T., Dou, Z., Sun, M., et al. (2025) · 2025
Closest in time.