Fetching the paper…
Reading the bibliography…
The integrity of AI benchmarks is fundamental to accurately assess the capabilities of AI systems.
The measurement of observer agreement for categorical data
JR Landis · 1977
Earlier work this paper cites.
Controlling the false discovery rate: a practical and powerful approach to multiple testing
Yoav Benjamini and Yosef Hochberg · 1995
Earlier work this paper cites.
Generalization of ‘same–different’classification abilities in bottlenosed dolphins
Eduardo Mercado III, Deirdre A Killebrew, Adam A Pack, Inés VB Mácha, and Louis M Herman · 2000
Earlier work this paper cites.
Aging and vocabulary score: A meta-analysis
Paul Verhaeghen · 2003
Earlier work this paper cites.
Fundamentals of clinical trials , volume 4
Lawrence M Friedman, Curt Furberg, David L DeMets, David M Reboussin, Christopher B Granger, et al · 2010
Earlier work this paper cites.
Interrater reliability: the kappa statistic
Mary L McHugh · 2012
Earlier work this paper cites.
Preliminary study of object labeling using sound production in a beluga
Murayama Tsukasa, Yuki Fujii, Takayuki Hashimoto, Aya Shimoda, So Iijima, Kohei Hayasaka, Narumi Shiroma, Mana Koshikawa, Hiroshi Katsumata, Makoto Soichi, et al · 2012
Earlier work this paper cites.
A simple method to determine if a music information retrieval system is a “horse”
Bob L Sturm · 2014
Earlier work this paper cites.
“why should I trust you?” explaining the predictions of any classifier
Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin · 2016
Earlier work this paper cites.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith · 2017
Earlier work this paper cites.
The argument reasoning comprehension task: Identification and reconstruction of implicit warrants
Ivan Habernal, Henning Wachsmuth, Iryna Gurevych, and Benno Stein · 2018
Earlier work this paper cites.
Hypothesis only baselines in natural language inference
Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme · 2018
Earlier work this paper cites.
Gazing into clever hans machines
Jose Hernandez-Orallo · 2019
Earlier work this paper cites.
When choosing plausible alternatives, Clever Hans can be clever
Pride Kavumba, Naoya Inoue, Benjamin Heinzerling, Keshav Singh, Paul Reisert, and Kentaro Inui · 2019
Earlier work this paper cites.
Adversarial NLI: A new benchmark for natural language understanding
Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela · 2019
Earlier work this paper cites.
Probing neural network comprehension of natural language arguments
Timothy Niven and Hung-Yu Kao · 2019
Earlier work this paper cites.
What does BERT learn from multiple-choice reading comprehension datasets?
Chenglei Si, Shuohang Wang, Min-Yen Kan, and Jing Jiang · 2019
Earlier work this paper cites.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei · 2020
Earlier work this paper cites.
Specification gaming: the flip side of AI ingenuity
Victoria Krakovna, Jonathan Uesato, Vladimir Mikulik, Matthew Rahtz, Tom Everitt, Ramana Kumar, Zac Kenton, Jan Leike, and Shane Legg · 2020
Earlier work this paper cites.
Assessing the benchmarking capacity of machine reading comprehension datasets
Saku Sugawara, Pontus Stenetorp, Kentaro Inui, and Akiko Aizawa · 2020
Earlier work this paper cites.
An empirical study on robustness to spurious correlations using pre-trained language models
Lifu Tu, Garima Lalwani, Spandana Gella, and He He · 2020
Earlier work this paper cites.
What will it take to fix benchmarking in natural language understanding?
Samuel R. Bowman and George E. Dahl · 2021
Cited alongside, same era.
Competency problems: On finding and removing artifacts in language data
Matt Gardner, William Merrill, Jesse Dodge, Matthew E Peters, Alexis Ross, Sameer Singh, and Noah A Smith · 2021
Cited alongside, same era.
Are we learning yet? a meta review of evaluation failures across machine learning
Thomas Liao, Rohan Taori, Deborah Raji, and Ludwig Schmidt · 2021
Cited alongside, same era.
Non-numerical strategies used by bees to solve numerical cognition tasks
HaDi MaBouDi, Andrew B Barron, Sun Li, Maria Honkanen, Olli J Loukola, Fei Peng, Wenfeng Li, James AR Marshall, Alex Cope, Eleni Vasilaki, et al · 2021
Cited alongside, same era.
CommonsenseQA 2.0: Exposing the limits of AI through gamification
CLadder: Assessing causal reasoning in language models
Zhijing Jin, Yuen Chen, Felix Leeb, Luigi Gresele, Ojasv Kamal, Zhiheng Lyu, Kevin Blin, Fernando Gonzalez, Max Kleiman-Weiner, Mrinmaya Sachan, and Bernhard Schölkopf · 2023
Later among the works it cites.
Man Luo, Shrinidhi Kumbhar, Mihir Parmar, Neeraj Varshney, Pratyay Banerjee, Somak Aditya, Chitta Baral, et al · 2023
Later among the works it cites.
Evaluating cognitive maps and planning in large language models with CogEval
Ida Momennejad, Hosein Hasanbeig, Felipe Vieira Frujeri, Hiteshi Sharma, Nebojsa Jojic, Hamid Palangi, Robert Ness, and Jonathan Larson · 2023
Later among the works it cites.
Automated annotation of meta-features for predicting language model performance in natural language processing tasks, 2023
Yael Moros Daval · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Alon Talmor, Ori Yoran, Ronan Le Bras, Chandrasekhar Bhagavatula, Yoav Goldberg, Yejin Choi, and Jonathan Berant · 2021
Cited alongside, same era.
On reality and the limits of language data, 2022
Nigel H. Collier, Fangyu Liu, and Ehsan Shareghi · 2022
Cited alongside, same era.
Finding dataset shortcuts with grammar induction
Dan Friedman, Alexander Wettig, and Danqi Chen · 2022
Cited alongside, same era.
The acoustic bases of human voice identity processing in dogs
Anna Gábor, Noémi Kaszás, Tamás Faragó, Paula Pérez Fraga, Melinda Lovas, and Attila Andics · 2022
Cited alongside, same era.
Are prompt-based models clueless?
Pride Kavumba, Ryo Takahashi, and Yusuke Oda · 2022
Cited alongside, same era.
Holistic evaluation of language models
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al · 2022
Cited alongside, same era.
Wanli: Worker and ai collaboration for natural language inference dataset creation
Alisa Liu, Swabha Swayamdipta, Noah A Smith, and Yejin Choi · 2022
Cited alongside, same era.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al · 2022
Cited alongside, same era.
Manley Roberts, Himanshu Thakur, Christine Herlihy, Colin White, and Samuel Dooley · 2023
Later among the works it cites.
Language models are greedy reasoners: A systematic formal analysis of chain-of-thought
Abulhair Saparov and He He · 2023
Later among the works it cites.
It takes two to tango: Navigating conceptualizations of NLP tasks and measurements of performance
Arjun Subramonian, Xingdi Yuan, Hal Daumé III, and Su Lin Blodgett · 2023
Later among the works it cites.
Glore: Evaluating logical reasoning of large language models
Zhiyang Teng, Ruoxi Ning, Jian Liu, Qiji Zhou, Yue Zhang, et al · 2023
Later among the works it cites.
A critical review of causal inference benchmarks for large language models
Linying Yang, Oscar Clivio, Vik Shirvaikar, and Fabian Falck · 2023
Later among the works it cites.
When benchmarks are targets: Revealing the sensitivity of large language model leaderboards
Norah Alzahrani, Hisham Abdullah Alyahya, Yazeed Alnumay, Sultan Alrashed, Shaykhah Alsubaie, Yusef Almushaykeh, Faisal Mirza, Nouf Alotaibi, Nora Altwairesh, Areeb Alowisheq, et al · 2024
Closest in time.
Lessons from the trenches on reproducible evaluation of language models
Stella Biderman, Hailey Schoelkopf, Lintang Sutawika, Leo Gao, Jonathan Tow, Baber Abbasi, Alham Fikri Aji, Pawan Sasanka Ammanamanchi, Sid Black, Jordan Clive, Anthony DiPofi, Julen Etxaniz, Benjamin Fattori, Jessica Zosa Forde, Charles Foster, Mimansa Jaiswal, Wilson Y. Lee, Haonan Li, Charles Lovering, Niklas Muennighoff, Ellie Pavlick, Jason Phang, Aviya Skowron, Samson Tan, Xiangru Tang, Kevin A. Wang, Genta Indra Winata, Franccois Yvon, and Andy Zou · 2024
Closest in time.
StructEval: Deepen and broaden large language model assessment via structured evaluation, 2024
Boxi Cao, Mengjie Ren, Hongyu Lin, Xianpei Han, Feng Zhang, Junfeng Zhan, and Le Sun · 2024
Closest in time.
Training on the test task confounds evaluation and emergence
Ricardo Dominguez-Olmedo, Florian E Dorner, and Moritz Hardt · 2024
Closest in time.
Under the radar?, 2024
Jones Elliot, Hardalupas Mahi, and Agnew William · 2024
Closest in time.
Auxiliary task demands mask the capabilities of smaller language models
Jennifer Hu and Michael C. Frank · 2024
Closest in time.
Does data contamination make a difference? insights from intentionally contaminating pre-training data for language models
Minhao Jiang, Ken Liu, Ming Zhong, Rylan Schaeffer, Siru Ouyang, Jiawei Han, and Sanmi Koyejo · 2024
Closest in time.
Inadequacies of large language model benchmarks in the era of generative artificial intelligence
Timothy R. McIntosh, Teo Susnjak, Tong Liu, Paul Watters, and Malka N. Halgamuge · 2024
Closest in time.
State of what art? a call for multi-prompt llm evaluation
Moran Mizrahi, Guy Kaplan, Dan Malkin, Rotem Dror, Dafna Shahaf, and Gabriel Stanovsky · 2024
Closest in time.
General interaction battery: Simple object navigation and affordances (GIBSONA)
Danaja Rutar, Lucy Gaia Cheke, José Hernández-Orallo, Alva Markelius, and Wout Schellaert · 2024
Closest in time.
Yuqing Wang and Yun Zhao · 2024
Closest in time.