Fetching the paper…
Reading the bibliography…
As Large Language Models (LLMs) and other AI systems evolve, robustly estimating their capabilities from inherently stochastic outputs while systematically quantifying uncertainty in these estimates becomes increasingly important.
“BoolQ: Exploring the Surprising Difficulty of Natural Yes/No Questions”, 2019
Christopher Clark et al · 1905
Earlier work this paper cites.
“Composable Effects for Flexible and Accelerated Probabilistic Programming in NumPyro”, 2019
Du Phan, Neeraj Pradhan and Martin Jankowiak · 1912
Earlier work this paper cites.
“Generalized Linear Models”
John Nelder and Robert William Wedderburn · 1972
Earlier work this paper cites.
“Measuring Massive Multitask Language Understanding”, 2020
Dan Hendrycks et al · 2009
Earlier work this paper cites.
“Asymptotic Equivalence of Bayes Cross Validation and Widely Applicable Information Criterion in Singular Learning Theory”
Sumio Watanabe · 2010
Earlier work this paper cites.
“The No-U-Turn Sampler: Adaptively Setting Path Lengths in Hamiltonian Monte Carlo”, 2011
Matthew. Hoffman and Andrew Gelman · 2011
Earlier work this paper cites.
“Bayesian Data Analysis”, 2013
Andrew Gelman et al · 2013
Earlier work this paper cites.
“RACE: Large-scale ReAding Comprehension Dataset From Examinations”, 2017
Guokun Lai et al · 2017
Earlier work this paper cites.
“Statistical Rethinking”
Richard McElreath · 2017
Earlier work this paper cites.
“Evaluating Large Language Models Trained on Code”, 2021
Mark Chen et al · 2021
Earlier work this paper cites.
“DS-1000: A Natural and Reliable Benchmark for Data Science Code Generation”, 2022
Yuhang Lai et al · 2022
Cited alongside, same era.
“Holistic Evaluation of Language Models”, 2022
Percy Liang et al · 2022
Cited alongside, same era.
“Language Models are Multilingual Chain-of-Thought Reasoners”, 2022
Freda Shi et al · 2022
Cited alongside, same era.
“ReAct: Synergizing Reasoning and Acting in Language Models”, 2022
Shunyu Yao et al · 2022
Cited alongside, same era.
“SWE-bench: Can Language Models Resolve Real-world GitHub Issues?”
Tianyu Han et al · 2023
Cited alongside, same era.
OpenAI et al · 2024
Later among the works it cites.
“Clio: Privacy-Preserving Insights into Real-World AI Use”, 2024
Alex Tamkin et al · 2024
Later among the works it cites.
“inspect_ai: AI-assisted inspection policy development”, https://github.com/UKGovernmentBEIS/inspect_ai , 2024
UKGovernmentBEIS · 2024
Later among the works it cites.
Zhaojian Yu, Yilun Zhao, Arman Cohan and Xiao-Ping Zhang · 2024
Later among the works it cites.
“HiBayES: A Python package for analysing data from Inspect logs using statistical modeling techniques.”, https://github.com/UKGovernmentBEIS/hibayes , 2025
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Statistical Rethinking 2023”, https://github.com/rmcelreath/stat_rethinking_2023 , 2023
Richard McElreath · 2023
Cited alongside, same era.
“GAIA: a benchmark for General AI Assistants”, 2023
Grégoire Mialon et al · 2023
Cited alongside, same era.
“InterCode: Standardizing and Benchmarking Interactive Coding with Execution Feedback”, 2023
John Yang, Akshara Prabhakar, Karthik Narasimhan and Shunyu Yao · 2023
Cited alongside, same era.
“The Claude 3 Model Family: Opus, Sonnet, Haiku Anthropic”, 2024
Anthropic · 2024
Cited alongside, same era.
“Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations”, 2024
Evan Miller · 2024
Cited alongside, same era.
Harry Coppock et al · 2025
Closest in time.
“Analyzing OpenAI’s o1-preview in Comparison with Claude 3 Opus and GPT-4 on ARC AGI Evaluation”, 2025
Greg Kamradt · 2025
Closest in time.
“Measuring AI Ability to Complete Long Tasks”, 2025
Thomas Kwa et al · 2025
Closest in time.
“METR’s GPT-4.5 pre-deployment evaluations”, https://metr.org/blog/2025-02-27-gpt-4-5-evals/ , 2025
METR · 2025
Closest in time.
“HCAST: Human-Calibrated Autonomy Software Tasks”, 2025
David Rein et al · 2025
Closest in time.