Fetching the paper…
Reading the bibliography…
Rigorous statistical evaluations of large language models (LLMs), including valid error bars and significance testing, are essential for meaningful and reliable performance assessment.
On the interpretation of χ \chi 2 from contingency tables, and the calculation of p
Fisher, R. A · 1922
Earlier work this paper cites.
Probable inference, the law of succession, and statistical inference
Wilson, E. B · 1927
Earlier work this paper cites.
The use of confidence or fiducial limits illustrated in the case of the binomial
Clopper, C. J. and Pearson, E. S · 1934
Earlier work this paper cites.
Objections to Bayesian statistics
Gelman, A · 1936
Earlier work this paper cites.
Exact bayesian analysis of a 2 × \times 2 contingency table, and fisher’s “exact” significance test
Altham, P. M · 1969
Earlier work this paper cites.
Why Isn’t Everyone a Bayesian?
Efron, B · 1986
Earlier work this paper cites.
A note on the delta method
Oehlert, G. W · 1992
Earlier work this paper cites.
Bayesian Interval Estimates which are also Confidence Intervals
Severini, T. A · 1993
Earlier work this paper cites.
Approximate is better than “exact” for interval estimation of binomial proportions
Agresti, A. and Coull, B. A · 1998
Earlier work this paper cites.
Interval estimation for the difference between independent proportions: comparison of eleven methods
Newcombe, R. G · 1998
Earlier work this paper cites.
Bootstrap tests: How many bootstraps?
Davidson, R. and MacKinnon, J. G · 2000
Earlier work this paper cites.
Asymptotic statistics , volume 3
Van der Vaart, A. W · 2000
Earlier work this paper cites.
Importance sampling: a review
Tokdar, S. T. and Kass, R. E · 2010
Earlier work this paper cites.
The widespread misinterpretation of p-values as error probabilities
Hubbard, R · 2011
Earlier work this paper cites.
In defence of score intervals for proportions and their differences
Newcombe, R. G. and Nurminen, M. M · 2011
Earlier work this paper cites.
The cost of using exact confidence intervals for a binomial proportion
Thulin, M · 2014
Earlier work this paper cites.
Statistical tests, p values, confidence intervals, and power: a guide to misinterpretations
Greenland, S., Senn, S. J., Rothman, K. J., Carlin, J. B., Poole, C., Goodman, S. N., and Altman, D. G · 2016
Cited alongside, same era.
A Bayesian interpretation of the confusion matrix
Caelen, O · 2017
Cited alongside, same era.
Race: Large-scale reading comprehension dataset from examinations
Lai, G., Xie, Q., Liu, H., Yang, Y., and Hovy, E · 2017
Cited alongside, same era.
Quac: Question answering in context
Choi, E., He, H., Iyyer, M., Yatskar, M., Yih, W.-t., Choi, Y., Liang, P., and Zettlemoyer, L · 2018
Cited alongside, same era.
The hitchhiker‘s guide to testing statistical significance in natural language processing
Dror, R., Baumer, G., Shlomov, S., and Reichart, R · 2018
Cited alongside, same era.
Know what you don‘t know: Unanswerable questions for SQuAD
Rajpurkar, P., Jia, R., and Liang, P · 2018
Lessons from the trenches on reproducible evaluation of language models
Biderman, S., Schoelkopf, H., Sutawika, L., Gao, L., Tow, J., Abbasi, B., Aji, A. F., Ammanamanchi, P. S., Black, S., Clive, J., et al · 2024
Later among the works it cites.
Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., Schelten, A., Yang, A., Fan, A., et al · 2024
Later among the works it cites.
Frontiermath: A benchmark for evaluating advanced mathematical reasoning in ai
Glazer, E., Erdil, E., Besiroglu, T., Chicharro, D., Chen, E., Gunning, A., Olsson, C. F., Denain, J.-S., Ho, A., Santos, E. d. O., et al · 2024
Later among the works it cites.
Experimental Design and Analysis for AI Researchers, 2024
Hermann, K., Hu, J., and Mozer, M · 2024
Later among the works it cites.
SWE-bench: Can language models resolve real-world github issues?
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Drop: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dua, D., Wang, Y., Dasigi, P., Stanovsky, G., Singh, S., and Gardner, M · 2019
Cited alongside, same era.
Training verifiers to solve math word problems
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., et al · 2021
Cited alongside, same era.
Scientific credibility of machine translation research: A meta-evaluation of 769 papers
Marie, B., Fujita, A., and Rubino, R · 2021
Cited alongside, same era.
AI and the everything in the whole wide world benchmark
Raji, I. D., Bender, E. M., Paullada, A., Denton, E., and Hanna, A · 2021
Cited alongside, same era.
A simple approximation for the bivariate normal integral
Tsay, W.-J. and Ke, P.-H · 2021
Cited alongside, same era.
Testing statistical hypotheses , volume 4
Lehmann, E. L., Romano, J. P., and Casella, G · 2022
Cited alongside, same era.
Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., and Narasimhan, K. R · 2024
Later among the works it cites.
Quantifying variance in evaluation benchmarks
Madaan, L., Singh, A. K., Schaeffer, R., Poulton, A., Koyejo, S., Stenetorp, P., Narang, S., and Hupkes, D · 2024
Later among the works it cites.
Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
Miller, E · 2024
Later among the works it cites.
Betterbench: Assessing ai benchmarks, uncovering issues, and establishing best practices
Reuel, A., Hardy, A., Smith, C., Lamparth, M., Hardy, M., and Kochenderfer, M. J · 2024
Later among the works it cites.
American invitational mathematics examination
AIME · 2025
Closest in time.
American invitational mathematics examination, 2025
AIME II · 2025
Closest in time.
Matharena: Evaluating llms on uncontaminated math competitions, February 2025
Balunović, M., Dekoninck, J., Petrov, I., Jovanović, N., and Vechev, M · 2025
Closest in time.
MLE-bench: Evaluating machine learning agents on machine learning engineering
Chan, J. S., Chowdhury, N., Jaffe, O., Aung, J., Sherburn, D., Mays, E., Starace, G., Liu, K., Maksin, L., Patwardhan, T., Madry, A., and Weng, L · 2025
Closest in time.
More than marketing? on the information value of ai benchmarks for practitioners
Hardy, A., Reuel, A., Jafari Meimandi, K., Soder, L., Griffith, A., Asmar, D. M., Koyejo, S., Bernstein, M. S., and Kochenderfer, M. J · 2025
Closest in time.
AI Models Are Getting Smarter. New Tests Are Racing to Catch Up, 2024
Pillay, T · 2025
Closest in time.
Livebench: A challenging, contamination-free LLM benchmark
White, C., Dooley, S., Roberts, M., Pal, A., Feuer, B., Jain, S., Shwartz-Ziv, R., Jain, N., Saifullah, K., Dey, S., Shubh-Agrawal, Sandha, S. S., Naidu, S. V., Hegde, C., LeCun, Y., Goldstein, T., Neiswanger, W., and Goldblum, M · 2025
Closest in time.