Fetching the paper…
Reading the bibliography…
Evaluations are critical for understanding the capabilities of large language models (LLMs).
“With Little Power Comes Great Responsibility”, 2020
Dallas Card et al · 2010
Earlier work this paper cites.
“So you want to run an experiment, now what? Some simple rules of thumb for optimal experimental design”, Working Paper Series 15701, 2010
John List, Sally Sadoff and Mathis Wagner · 2010
Earlier work this paper cites.
“Distilling the Knowledge in a Neural Network”, 2015
Geoffrey Hinton, Oriol Vinyals and Jeff Dean · 2015
Earlier work this paper cites.
“Causal Inference for Statistics, Social, and Biomedical Sciences: An Introduction”
Guido. Imbens and Donald. Rubin · 2015
Earlier work this paper cites.
“RACE: Large-scale ReAding Comprehension Dataset From Examinations”
Guokun Lai et al · 2017
Earlier work this paper cites.
“QuAC : Question Answering in Context”, 2018
Eunsol Choi et al · 2018
Earlier work this paper cites.
“Know What You Don’t Know: Unanswerable Questions for SQuAD”, 2018
Pranav Rajpurkar, Robin Jia and Percy Liang · 2018
Cited alongside, same era.
“DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs”
Dheeru Dua et al · 2019
Cited alongside, same era.
“Evaluating Large Language Models Trained on Code”, 2021
Mark Chen et al · 2021
Cited alongside, same era.
“Measuring Mathematical Problem Solving With the MATH Dataset”, 2021
Dan Hendrycks et al · 2021
Cited alongside, same era.
“When should you adjust standard errors for clustering?”
Alberto Abadie, Susan Athey, Guido Imbens and Jeffrey Wooldridge · 2022
Cited alongside, same era.
“Challenges in evaluating AI systems”, 2023
Deep Ganguli, Nicholas Schiefer, Marina Favaro and Jack Clark · 2023
Later among the works it cites.
“Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference”, 2024
Wei-Lin Chiang et al · 2024
Closest in time.
“The Llama 3 Herd of Models”, 2024
Abhimanyu Dubey et al · 2024
Closest in time.
“Inspect: An open-source framework for large language model evaluations” Accessed: 2024-09-03, https://inspect.ai-safety-institute.org.uk/
2024
Closest in time.
“Quantifying Variance in Evaluation Benchmarks”, 2024
Lovish Madaan et al · 2024
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
“Language Models are Multilingual Chain-of-Thought Reasoners”, 2022
Freda Shi et al · 2022
Cited alongside, same era.
“OpenAI Evals” Accessed: 2024-09-03, https://github.com/openai/evals
2024
Closest in time.