Fetching the paper…
Reading the bibliography…
Multiple choice question answering (MCQA) is popular for LLM evaluation due to its simplicity and human-like testing, but we argue for its reform.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
The kansas silent reading tests
Frederick James Kelly. 1916 · 1916
Earlier work this paper cites.
A report on the use of the kansas silent reading tests with over one hundred thousand children
Walter S Monroe. 1917 · 1917
Earlier work this paper cites.
Psychology in relation to the war
Robert M Yerkes. 1918 · 1918
Earlier work this paper cites.
Methods for discriminating levels of partial knowledge concerning a test item
Bruno de Finetti. 1965 · 1965
Earlier work this paper cites.
The college board admissions testing program: A technical report on research and development activities relating to the scholastic aptitude test and achievement tests
William H Angoff. 1971 · 1971
Earlier work this paper cites.
The IBM Watson laboratory at Columbia university: a history
Jean Ford Brennan and HK Clark. 1971 · 1971
Earlier work this paper cites.
The validity of using holistic scoring to evaluate writing: A critical overview
Davida Charney. 1984 · 1984
Earlier work this paper cites.
Foundations and grand challenges of artificial intelligence: Aaai presidential address
Raj Reddy. 1988 · 1988
Earlier work this paper cites.
A taxonomy of multiple-choice item-writing rules
Thomas M Haladyna and Steven M Downing. 1989 · 1989
Earlier work this paper cites.
A description of what happens when an examinee takes a multiple-choice reading comprehension test
Roger Farr, Robert Pritchard, and Brian Smitten. 1990 · 1990
Earlier work this paper cites.
Cognitive complexity and the comparability of multiple-choice and constructed-response test formats
Gregory R Hancock. 1994 · 1994
Earlier work this paper cites.
A comparative study of measures of partial knowledge in multiple-choice tests
Anat Ben-Simon, David V Budescu, and Baruch Nevo. 1997 · 1997
Earlier work this paper cites.
Toefl 2000 framework
Joan Jamieson, Stan Jones, Irwin Kirsch, Peter Mosenthal, and Carol Taylor. 2000 · 2000
Earlier work this paper cites.
The basics of item response theory
Frank B Baker. 2001 · 2001
Earlier work this paper cites.
Using working memory theory to investigate the construct validity of multiple-choice reading comprehension tests such as the sat
Meredyth Daneman and Brenda Hannon. 2001 · 2001
Earlier work this paper cites.
Writing multiple-choice test items that promote and measure critical thinking
Susan Morrison and Kathleen Walsh Free. 2001 · 2001
Earlier work this paper cites.
Multiple choice questions: their value as an assessment tool
Edward Moss. 2001 · 2001
Earlier work this paper cites.
Gre scores as predictors of minority students’success in graduate study: An argument for change
Charles Sampson and Patricia G Boyer. 2001 · 2001
Earlier work this paper cites.
The trec question answering track
Ellen M Voorhees. 2001 · 2001
Earlier work this paper cites.
Columbia university professor ben wood
Ben Wood and Reynold Johnson. 2001 · 2001
Earlier work this paper cites.
A review of multiple-choice item-writing guidelines for classroom assessment
Thomas M Haladyna, Steven M Downing, and Michael C Rodriguez. 2002 · 2002
Earlier work this paper cites.
A revision bloom’s taxonomy: An overview
DR Krathwohl. 2002 · 2002
Earlier work this paper cites.
Combining independent modules to solve multiple-choice synonym and analogy problems
Peter D Turney, Michael L Littman, Jeffrey Bigham, and Victor Shnayder. 2003 · 2003
Earlier work this paper cites.
Wordnet sits the sat a knowledge-based approach to lexical analogy
Tony Veale. 2004 · 2004
Earlier work this paper cites.
Multiple-choice tests and student understanding: What is the connection?
Mark G Simkin and William L Kuechler. 2005 · 2005
Earlier work this paper cites.
An analysis of negative marking in multiple-choice assessment
Alan Holt. 2006 · 2006
Earlier work this paper cites.
18 multidimensional item response theory
Mark D Reckase. 2006 · 2006
Earlier work this paper cites.
Multiple choice questions: Can they examine application of knowledge?
Ieva Stupans. 2006 · 2006
Earlier work this paper cites.
Brainiac: adventures in the curious, competitive, compulsive world of trivia buffs
Ken Jennings. 2007 · 2007
Earlier work this paper cites.
Assessment of higher order cognitive skills in undergraduate education: modified essay or multiple choice questions? research paper
Edward J Palmer and Peter G Devitt. 2007 · 2007
Earlier work this paper cites.
Statistical theories of mental test scores
Frederic M Lord and Melvin R Novick. 2008 · 2008
Earlier work this paper cites.
Open book testing in online learning environments
Glenda C Rakes. 2008 · 2008
Earlier work this paper cites.
Constructed-response test questions: Why we use them; how we score them. r&d connections. number 11
Samuel A Livingston. 2009 · 2009
Earlier work this paper cites.
Arguing to learn and learning to argue: Design justifications and guidelines
David H Jonassen and Bosung Kim. 2010 · 2010
Earlier work this paper cites.
Improving employees’ compliance through information systems security training: an action research study
Petri Puhakainen and Mikko Siponen. 2010 · 2010
Earlier work this paper cites.
How to write good multiple-choice questions
Dianne E Campbell. 2011 · 2011
Earlier work this paper cites.
Guessing, partial knowledge, and misconceptions in multiple-choice tests
Paul Ngee Kiong Lau, Sie Hoe Lau, Kian Sam Hong, and Hasbee Usop. 2011 · 2011
Earlier work this paper cites.
Validating measurement of knowledge integration in science using multiple-choice and explanation items
Hee-Sun Lee, Ou Lydia Liu, and Marcia C Linn. 2011 · 2011
Earlier work this paper cites.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S Gordon. 2011 · 2011
Earlier work this paper cites.
The winograd schema challenge
Hector Levesque, Ernest Davis, and Leora Morgenstern. 2012 · 2012
Earlier work this paper cites.
Can multiple-choice questions simulate free-response questions?
Shih-Yin Lin and Chandralekha Singh. 2012 · 2012
Earlier work this paper cites.
Construct validity and constructed-response tests
Richard E Snow. 2012 · 2012
Earlier work this paper cites.
Is there a case for driver training? a review of the efficacy of pre-and post-licence driver training
Vanessa Beanland, Natassia Goode, Paul M Salmon, and Michael G Lenné. 2013 · 2013
Earlier work this paper cites.
A study of the knowledge base requirements for passing an elementary science test
Peter Clark, Philip Harrison, and Niranjan Balasubramanian. 2013 · 2013
Earlier work this paper cites.
Comparing comprehension measured by multiple-choice and open-ended questions
Yasuhiro Ozuru, Stephen Briner, Christopher A Kurby, and Danielle S McNamara. 2013 · 2013
Earlier work this paper cites.
Mctest: A challenge dataset for the open-domain machine comprehension of text
Matthew Richardson, Christopher JC Burges, and Erin Renshaw. 2013 · 2013
Earlier work this paper cites.
A comparison of learning environments: All that glitters…
Valerie J Shute. 2013 · 2013
Earlier work this paper cites.
Can an ai get into the university of tokyo?
Eliza Strickland. 2013 · 2013
Earlier work this paper cites.
Distractor quality evaluation in multiple choice questions
Van-Minh Pho, Anne-Laure Ligozat, and Brigitte Grau. 2015 · 2015
Earlier work this paper cites.
Snapshots of mathematics teacher noticing during task design
Ban Heng Choy. 2016 · 2016
Earlier work this paper cites.
From extractive to abstractive summarization: A journey
Parth Mehta. 2016 · 2016
Earlier work this paper cites.
Developing, analyzing, and using distractors for multiple-choice tests in education: A comprehensive review
Mark J Gierl, Okan Bulut, Qi Guo, and Xinxin Zhang. 2017 · 2017
Earlier work this paper cites.
Making the v in vqa matter: Elevating the role of image understanding in visual question answering
Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. 2017 · 2017
Earlier work this paper cites.
On calibration of modern neural networks
Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q Weinberger. 2017 · 2017
Earlier work this paper cites.
RACE: Large-scale ReAding comprehension dataset from examinations
Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. 2017 · 2017
Earlier work this paper cites.
The limitations of the gre in predicting success in biomedical graduate school
Liane Moneta-Koehler, Abigail M Brown, Kimberly A Petrie, Brent J Evans, and Roger Chalkley. 2017 · 2017
Earlier work this paper cites.
Large language models sensitivity to the order of options in multiple-choice questions
Pouya Pezeshkpour and Estevam Hruschka. 2024 · 2017
Earlier work this paper cites.
Multiple-choice testing in education: Are the best practices for assessment also good for learning?
Andrew C Butler. 2018 · 2018
Earlier work this paper cites.
Think you have solved question answering? try arc, the ai2 reasoning challenge
Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018 · 2018
Earlier work this paper cites.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. 2018 · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018 · 2018
Earlier work this paper cites.
What makes reading comprehension questions easier?
Saku Sugawara, Kentaro Inui, Satoshi Sekine, and Akiko Aizawa. 2018 · 2018
Earlier work this paper cites.
Elimination testing with adapted scoring reduces guessing and anxiety in multiple-choice assessments, but does not increase grade average in comparison with negative marking
Jef Vanderoost, Rianne Janssen, Jan Eggermont, Riet Callens, and Tinne De Laet. 2018 · 2018
Earlier work this paper cites.
SWAG: A large-scale adversarial dataset for grounded commonsense inference
Rowan Zellers, Yonatan Bisk, Roy Schwartz, and Yejin Choi. 2018 · 2018
Earlier work this paper cites.
Understanding dataset design choices for multi-hop reasoning
Jifan Chen and Greg Durrett. 2019 · 2019
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Dheeru Dua, Yizhong Wang, Pradeep Dasigi, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019 · 2019
Earlier work this paper cites.
Abstractive summarization: A survey of the state of the art
Hui Lin and Vincent Ng. 2019 · 2019
Earlier work this paper cites.
Multiple choice questions: answering correctly and knowing the answer
Peter McKenna. 2019 · 2019
Earlier work this paper cites.
Social IQa: Commonsense reasoning about social interactions
Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, and Yejin Choi. 2019 · 2019
Earlier work this paper cites.
CommonsenseQA: A question answering challenge targeting commonsense knowledge
Alon Talmor, Jonathan Herzig, Nicholas Lourie, and Jonathan Berant. 2019 · 2019
Earlier work this paper cites.
Trick me if you can: Human-in-the-loop generation of adversarial examples for question answering
Eric Wallace, Pedro Rodriguez, Shi Feng, Ikuya Yamada, and Jordan Boyd-Graber. 2019 · 2019
Earlier work this paper cites.
HellaSwag: Can a machine really finish your sentence?
Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019 · 2019
Earlier work this paper cites.
Piqa: Reasoning about physical commonsense in natural language
Yonatan Bisk, Rowan Zellers, Jianfeng Gao, Yejin Choi, et al. 2020 · 2020
Cited alongside, same era.
What question answering can learn from trivia nerds
Jordan Boyd-Graber and Benjamin Börschinger. 2020 · 2020
Cited alongside, same era.
Evaluating models’ local decision boundaries via contrast sets
Matt Gardner, Yoav Artzi, Jonathan Berant, Ben Bogin, Sihao Chen, Dheeru Dua, Yanai Elazar, Ananth Gottumukkala, Nitish Gupta, Hannaneh Hajishirzi, Gabriel Ilharco, Daniel Khashabi, Kevin Lin, Jiangming Liu, Nelson F. Liu, Phoebe Mulcaire, Qiang Ning, Sameer Singh, Noah A. Smith, Sanjay Subramanian, Eric Wallace, Ally Zhang, and Ben Zhou. 2020 · 2020
Cited alongside, same era.
Learning the difference that makes a difference with counterfactually-augmented data
Divyansh Kaushik, Eduard Hovy, and Zachary Lipton. 2020 · 2020
Cited alongside, same era.
Linguistically-informed transformations (LIT): A method for automatically generating contrast sets
Chuanrong Li, Lin Shengshuo, Zeyu Liu, Xinyi Wu, Xuhui Zhou, and Shane Steinert-Threlkeld. 2020 · 2020
Cited alongside, same era.
Chatbot arena: An open platform for evaluating llms by human preference
Wei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos, Tianle Li, Dacheng Li, Hao Zhang, Banghua Zhu, Michael Jordan, Joseph E. Gonzalez, and Ion Stoica. 2024 · 2024
Later among the works it cites.
Fairness in large language models: A taxonomic survey
Zhibo Chu, Zichong Wang, and Wenbin Zhang. 2024 · 2024
Later among the works it cites.
Ticking all the boxes: Generated checklists improve llm evaluation and generation
Jonathan Cook, Tim Rocktäschel, Jakob Foerster, Dennis Aumiller, and Alex Wang. 2024 · 2024
Later among the works it cites.
Measuring and improving attentiveness to partial inputs with counterfactuals
Yanai Elazar, Bhargavi Paranjape, Hao Peng, Sarah Wiegreffe, Khyathi Chandu, Vivek Srikumar, Sameer Singh, and Noah A. Smith. 2024 · 2024
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The prisma 2020 statement: an updated guideline for reporting systematic reviews
Matthew J Page, Joanne E McKenzie, Patrick M Bossuyt, Isabelle Boutron, Tammy C Hoffmann, Cynthia D Mulrow, Larissa Shamseer, Jennifer M Tetzlaff, Elie A Akl, Sue E Brennan, et al. 2021 · 2020
Cited alongside, same era.
COMET: A neural framework for MT evaluation
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020 · 2020
Cited alongside, same era.
What does my qa model know? devising controlled probes using expert knowledge
Kyle Richardson and Ashish Sabharwal. 2020 · 2020
Cited alongside, same era.
Getting closer to ai complete question answering: A set of prerequisite real tasks
Anna Rogers, Olga Kovaleva, Matthew Downey, and Anna Rumshisky. 2020 · 2020
Cited alongside, same era.
Enhancing automated essay scoring performance via fine-tuning pre-trained language models with combination of regression and ranking
Ruosong Yang, Jiannong Cao, Zhiyuan Wen, Youzheng Wu, and Xiaodong He. 2020 · 2020
Cited alongside, same era.
On the application of transformers for estimating the difficulty of multiple-choice questions from text
Luca Benedetto, Giovanni Aradelli, Paolo Cremonesi, Andrea Cappelli, Andrea Giussani, and Roberto Turrin. 2021 · 2021
Cited alongside, same era.
Sumithra Bhakthavatsalam, Daniel Khashabi, Tushar Khot, Bhavana Dalvi Mishra, Kyle Richardson, Ashish Sabharwal, Carissa Schoenick, Oyvind Tafjord, and Peter Clark. 2021 · 2021
Cited alongside, same era.
Meng Fang, Xiangpeng Wan, Fei Lu, Fei Xing, and Kai Zou. 2024 · 2024
Later among the works it cites.
Exploring automated distractor generation for math multiple-choice questions via large language models
Wanyong Feng, Jaewook Lee, Hunter McNichols, Alexander Scarlatos, Digory Smith, Simon Woodhead, Nancy Ornelas, and Andrew Lan. 2024 · 2024
Later among the works it cites.
Open llm leaderboard v2
Clémentine Fourrier, Nathan Habib, Alina Lozovskaya, Konrad Szafer, and Thomas Wolf. 2024 · 2024
Later among the works it cites.
Do great minds think alike? investigating human-AI complementarity in question answering with CAIMIRA
Maharshi Gor, Hal Daumé Iii, Tianyi Zhou, and Jordan Lee Boyd-Graber. 2024 · 2024
Later among the works it cites.
Wait, that’s not an option: Llms robustness with incorrect multiple-choice options
Gracjan Góral, Emilia Wiśnios, Piotr Sankowski, and Paweł Budzianowski. 2024 · 2024
Later among the works it cites.
Alignment faking in large language models
Ryan Greenblatt, Carson Denison, Benjamin Wright, Fabien Roger, Monte MacDiarmid, Sam Marks, Johannes Treutlein, Tim Belonax, Jack Chen, David Duvenaud, et al. 2024 · 2024
Later among the works it cites.
Jiawei Gu, Xuhui Jiang, Zhichao Shi, Hexiang Tan, Xuehao Zhai, Chengjin Xu, Wei Li, Yinghan Shen, Shengjie Ma, Honghao Liu, et al. 2024 · 2024
Later among the works it cites.
Aaron Jaech, Adam Kalai, Adam Lerer, Adam Richardson, Ahmed El-Kishky, Aiden Low, Alec Helyar, Aleksander Madry, Alex Beutel, Alex Carney, et al. 2024 · 2024
Later among the works it cites.
A study on large language models’ limitations in multiple-choice question answering
Aisha Khatun and Daniel G Brown. 2024 · 2024
Later among the works it cites.
Kyungha Kim, Sangyun Lee, Kung-Hsiang Huang, Hou Pong Chan, Manling Li, and Heng Ji. 2024 · 2024
Later among the works it cites.
A systematic survey and critical review on evaluating large language models: Challenges, limitations, and recommendations
Md Tahmid Rahman Laskar, Sawsan Alqahtani, M Saiful Bari, Mizanur Rahman, Mohammad Abdullah Matin Khan, Haidar Khan, Israt Jahan, Amran Bhuiyan, Chee Wei Tan, Md Rizwan Parvez, Enamul Hoque, Shafiq Joty, and Jimmy Huang. 2024 · 2024
Later among the works it cites.
This land is Your, My land: Evaluating geopolitical bias in language models through territorial disputes
Bryan Li, Samar Haider, and Chris Callison-Burch. 2024a · 2024
Later among the works it cites.
CMMLU: Measuring massive multitask language understanding in Chinese
Haonan Li, Yixuan Zhang, Fajri Koto, Yifei Yang, Hai Zhao, Yeyun Gong, Nan Duan, and Timothy Baldwin. 2024b · 2024
Later among the works it cites.
Think twice before trusting: Self-detection for large language models through comprehensive answer reflection
Moxin Li, Wenjie Wang, Fuli Feng, Fengbin Zhu, Qifan Wang, and Tat-Seng Chua. 2024c · 2024
Later among the works it cites.
Anchored answers: Unravelling positional bias in gpt-2’s multiple-choice questions
Ruizhe Li and Yanjun Gao. 2024 · 2024
Later among the works it cites.
Can multiple-choice questions really be useful in detecting the abilities of LLMs?
Wangyue Li, Liangzhi Li, Tong Xiang, Xiao Liu, Wei Deng, and Noa Garcia. 2024e · 2024
Later among the works it cites.
PEDANTS: Cheap but effective and interpretable answer equivalence
Zongxia Li, Ishani Mondal, Huy Nghiem, Yijun Liang, and Jordan Lee Boyd-Graber. 2024g · 2024
Later among the works it cites.
Yinhong Liu, Zhijiang Guo, Tianya Liang, Ehsan Shareghi, Ivan Vulić, and Nigel Collier. 2024 · 2024
Later among the works it cites.
Beyond probabilities: Unveiling the misalignment in evaluating large language models
Chenyang Lyu, Minghao Wu, and Alham Aji. 2024 · 2024
Later among the works it cites.
Are self-explanations from large language models faithful?
Andreas Madsen, Sarath Chandar, and Siva Reddy. 2024 · 2024
Later among the works it cites.
Hybrid preferences: Learning to route instances for human vs. ai feedback
Lester James V Miranda, Yizhong Wang, Yanai Elazar, Sachin Kumar, Valentina Pyatkin, Faeze Brahman, Noah A Smith, Hannaneh Hajishirzi, and Pradeep Dasigi. 2024 · 2024
Later among the works it cites.
Characterizing large language models as rationalizers of knowledge-intensive tasks
Aditi Mishra, Sajjadur Rahman, Kushan Mitra, Hannah Kim, and Estevam Hruschka. 2024 · 2024
Later among the works it cites.
An automatic question usability evaluation toolkit
Steven Moore, Eamon Costello, Huy A Nguyen, and John Stamper. 2024 · 2024
Later among the works it cites.
Aidar Myrzakhan, Sondos Mahmoud Bsharat, and Zhiqiang Shen. 2024 · 2024
Later among the works it cites.
BLEnd: A benchmark for LLMs on everyday knowledge in diverse cultures and languages
Junho Myung, Nayeon Lee, Yi Zhou, Jiho Jin, Rifki Afina Putri, Dimosthenis Antypas, Hsuvas Borkakoty, Eunsu Kim, Carla Perez-Almendros, Abinew Ali Ayele, Victor Gutierrez Basulto, Yazmin Ibanez-Garcia, Hwaran Lee, Shamsuddeen Hassan Muhammad, Kiwoong Park, Anar Sabuhi Rzayev, Nina White, Seid Muhie Yimam, Mohammad Taher Pilehvar, Nedjma Ousidhoum, Jose Camacho-Collados, and Alice Oh. 2024 · 2024
Later among the works it cites.
Leaving the barn door open for clever hans: Simple features predict llm benchmark answers
Lorenzo Pacchiardi, Marko Tesic, Lucy G Cheke, and José Hernández-Orallo. 2024 · 2024
Later among the works it cites.
Plausibly problematic questions in multiple-choice benchmarks for commonsense reasoning
Shramay Palta, Nishant Balepur, Peter Rankel, Sarah Wiegreffe, Marine Carpuat, and Rachel Rudinger. 2024 · 2024
Later among the works it cites.
Making reasoning matter: Measuring and improving faithfulness of chain-of-thought reasoning
Debjit Paul, Robert West, Antoine Bosselut, and Boi Faltings. 2024 · 2024
Later among the works it cites.
Efficient benchmarking (of language models)
Yotam Perlitz, Elron Bandel, Ariel Gera, Ofir Arviv, Liat Ein-Dor, Eyal Shnarch, Noam Slonim, Michal Shmueli-Scheuer, and Leshem Choshen. 2024 · 2024
Later among the works it cites.
tinybenchmarks: evaluating llms with fewer examples
Felipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun, Gongjun Xu, and Mikhail Yurochkin. 2024 · 2024
Later among the works it cites.
GPQA: A graduate-level google-proof q&a benchmark
David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R. Bowman. 2024 · 2024
Later among the works it cites.
Benchmarks as microscopes: A call for model metrology
Michael Saxon, Ari Holtzman, Peter West, William Yang Wang, and Naomi Saphra. 2024 · 2024
Later among the works it cites.
The prompt report: A systematic survey of prompting techniques
Sander Schulhoff, Michael Ilie, Nishant Balepur, Konstantine Kahadze, Amanda Liu, Chenglei Si, Yinheng Li, Aayush Gupta, HyoJung Han, Sevien Schulhoff, et al. 2024 · 2024
Later among the works it cites.
KoCommonGEN v2: A benchmark for navigating Korean commonsense reasoning challenges in large language models
Jaehyung Seo, Jaewook Lee, Chanjun Park, SeongTae Hong, Seungjun Lee, and Heuiseok Lim. 2024 · 2024
Later among the works it cites.
Who validates the validators? aligning llm-assisted evaluation of llm outputs with human preferences
Shreya Shankar, JD Zamfirescu-Pereira, Björn Hartmann, Aditya Parameswaran, and Ian Arawjo. 2024 · 2024
Later among the works it cites.
From generation to selection: Findings of converting analogical problem-solving into multiple-choice questions
Donghyeon Shin, Seungpil Lee, Klea Lena Kovacec, and Sundong Kim. 2024 · 2024
Later among the works it cites.
Large language models help humans verify truthfulness – except when they are convincingly wrong
Chenglei Si, Navita Goyal, Tongshuang Wu, Chen Zhao, Shi Feng, Hal Daumé Iii, and Jordan Boyd-Graber. 2024 · 2024
Later among the works it cites.
Global mmlu: Understanding and addressing cultural and linguistic biases in multilingual evaluation
Shivalika Singh, Angelika Romanou, Clémentine Fourrier, David I Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, et al. 2024 · 2024
Later among the works it cites.
EDEN: Empathetic dialogues for English learning
Li Siyan, Teresa Shao, Zhou Yu, and Julia Hirschberg. 2024 · 2024
Later among the works it cites.
Dolma: an open corpus of three trillion tokens for language model pretraining research
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Evan Walsh, Luke Zettlemoyer, Noah Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, and Kyle Lo. 2024 · 2024
Later among the works it cites.
Polina Tsvilodub, Hening Wang, Sharon Grosch, and Michael Franke. 2024 · 2024
Later among the works it cites.
Language models don’t always say what they think: unfaithful explanations in chain-of-thought prompting
Miles Turpin, Julian Michael, Ethan Perez, and Samuel Bowman. 2024 · 2024
Later among the works it cites.
Investigating and addressing hallucinations of llms in tasks involving negation
Neeraj Varshney, Satyam Raj, Venkatesh Mishra, Agneet Chatterjee, Ritika Sarkar, Amir Saeidi, and Chitta Baral. 2024 · 2024
Later among the works it cites.
Unveiling selection biases: Exploring order and token sensitivity in large language models
Sheng-Lun Wei, Cheng-Kuang Wu, Hen-Hsen Huang, and Hsin-Hsi Chen. 2024 · 2024
Later among the works it cites.
Answer, assemble, ace: Understanding how transformers answer multiple choice questions
Sarah Wiegreffe, Oyvind Tafjord, Yonatan Belinkov, Hanna Hajishirzi, and Ashish Sabharwal. 2024 · 2024
Later among the works it cites.
Top leaderboard ranking = top coding proficiency, always? evoeval: Evolving coding benchmarks via LLM
Chunqiu Steven Xia, Yinlin Deng, and LINGMING ZHANG. 2024 · 2024
Later among the works it cites.
Strengthened symbol binding makes large language models reliable multiple-choice selectors
Mengge Xue, Zhenyu Hu, Liqun Liu, Kuo Liao, Shuang Li, Honglin Han, Meng Zhao, and Chengguo Yin. 2024 · 2024
Later among the works it cites.
CMoralEval: A moral evaluation benchmark for Chinese large language models
Linhao Yu, Yongqi Leng, Yufei Huang, Shang Wu, Haixin Liu, Xinmeng Ji, Jiahui Zhao, Jinwang Song, Tingting Cui, Xiaoqing Cheng, Liutao Liutao, and Deyi Xiong. 2024b · 2024
Later among the works it cites.
Wildchat: 1m chatGPT interaction logs in the wild
Wenting Zhao, Xiang Ren, Jack Hessel, Claire Cardie, Yejin Choi, and Yuntian Deng. 2024 · 2024
Later among the works it cites.
Revisiting the self-consistency challenges in multi-choice question formats for large language model evaluation
Wenjie Zhou, Qiang Wang, Mingzhou Xu, Ming Chen, and Xiangyu Duan. 2024a · 2024
Later among the works it cites.
Fool your (vision and) language model with embarrassingly simple permutations
Yongshuo Zong, Tingyang Yu, Ruchika Chavhan, Bingchen Zhao, and Timothy M. Hospedales. 2024 · 2024
Later among the works it cites.
Phd knowledge not required: A reasoning challenge for large language models
Carolyn Jane Anderson, Joydeep Biswas, Aleksander Boruch-Gruszecki, Federico Cassano, Molly Q Feldman, Arjun Guha, Francesca Lucchetti, and Zixuan Wu. 2025 · 2025
Closest in time.
ProverbEval: Exploring LLM evaluation challenges for low-resource language understanding
Israel Abebe Azime, Atnafu Lambebo Tonja, Tadesse Destaw Belay, Yonas Chanie, Bontu Fufa Balcha, Negasi Haile Abadi, Henok Biadglign Ademtew, Mulubrhan Abebe Nerea, Debela Desalegn Yadeta, Derartu Dagne Geremew, Assefa Atsbiha Tesfu, Philipp Slusallek, Thamar Solorio, and Dietrich Klakow. 2025 · 2025
Closest in time.
Reverse question answering: Can an LLM write a question so hard (or bad) that it can‘t answer?
Nishant Balepur, Feng Gu, Abhilasha Ravichander, Shi Feng, Jordan Lee Boyd-Graber, and Rachel Rudinger. 2025 · 2025
Closest in time.
Are we done with MMLU?
Aryo Pradipta Gema, Joshua Ong Jun Leang, Giwon Hong, Alessio Devoto, Alberto Carlo Maria Mancino, Rohit Saxena, Xuanli He, Yu Zhao, Xiaotang Du, Mohammad Reza Ghasemi Madani, Claire Barale, Robert McHardy, Joshua Harris, Jean Kaddour, Emile Van Krieken, and Pasquale Minervini. 2025 · 2025
Closest in time.
OLMES: A standard for language model evaluations
Yuling Gu, Oyvind Tafjord, Bailey Kuehl, Dany Haddad, Jesse Dodge, and Hannaneh Hajishirzi. 2025 · 2025
Closest in time.
Improving model evaluation using SMART filtering of benchmark datasets
Vipul Gupta, Candace Ross, David Pantoja, Rebecca J. Passonneau, Megan Ung, and Adina Williams. 2025b · 2025
Closest in time.
The BiGGen bench: A principled benchmark for fine-grained evaluation of language models with language models
Seungone Kim, Juyoung Suk, Ji Yong Cho, Shayne Longpre, Chaeeun Kim, Dongkeun Yoon, Guijin Son, Yejin Cho, Sheikh Shafayat, Jinheon Baek, Sue Hyun Park, Hyeonbin Hwang, Jinkyung Jo, Hyowon Cho, Haebin Shin, Seongyun Lee, Hanseok Oh, Noah Lee, Namgyu Ho, Se June Joo, Miyoung Ko, Yoonjoo Lee, Hyungjoo Chae, Jamin Shin, Joel Jang, Seonghyeon Ye, Bill Yuchen Lin, Sean Welleck, Graham Neubig, Moontae Lee, Kyungjae Lee, and Minjoon Seo. 2025 · 2025
Closest in time.
Wildbench: Benchmarking LLMs with challenging tasks from real users in the wild
Bill Yuchen Lin, Yuntian Deng, Khyathi Chandu, Abhilasha Ravichander, Valentina Pyatkin, Nouha Dziri, Ronan Le Bras, and Yejin Choi. 2025 · 2025
Closest in time.
DeLLMa: Decision making under uncertainty with large language models
Ollie Liu, Deqing Fu, Dani Yogatama, and Willie Neiswanger. 2025 · 2025
Closest in time.
LLMs are biased towards output formats! systematically evaluating and mitigating output format bias of LLMs
Do Xuan Long, Ngoc-Hai Nguyen, Tiviatis Sim, Hieu Dao, Shafiq Joty, Kenji Kawaguchi, Nancy F. Chen, and Min-Yen Kan. 2025 · 2025
Closest in time.
The realhumaneval: Evaluating large language models’ abilities to support programmers
Hussein Mozannar, Valerie Chen, Mohammed Alsobay, Subhro Das, Sebastian Zhao, Dennis Wei, Manish Nagireddy, Prasanna Sattigeri, Ameet Talwalkar, and David Sontag. 2025 · 2025
Closest in time.
Long Phan, Alice Gatti, Ziwen Han, Nathaniel Li, Josephina Hu, Hugh Zhang, Sean Shi, Michael Choi, Anish Agrawal, Arnav Chopra, et al. 2025 · 2025
Closest in time.
Who does the giant number pile like best: Analyzing fairness in hiring contexts
Preethi Seshadri and Seraphina Goldfarb-Tarrant. 2025 · 2025
Closest in time.
KMMLU: Measuring massive multitask language understanding in Korean
Guijin Son, Hanwool Lee, Sungdong Kim, Seungone Kim, Niklas Muennighoff, Taekyoon Choi, Cheonbok Park, Kang Min Yoo, and Stella Biderman. 2025 · 2025
Closest in time.
Is your benchmark truly adversarial? AdvScore: Evaluating human-grounded adversarialness
Yoo Yeon Sung, Maharshi Gor, Eve Fleisig, Ishani Mondal, and Jordan Lee Boyd-Graber. 2025 · 2025
Closest in time.
Llms may perform mcqa by selecting the least incorrect option
Haochun Wang, Sendong Zhao, Zewen Qiang, Nuwa Xi, Bing Qin, and Ting Liu. 2025 · 2025
Closest in time.
Language models learn to mislead humans via RLHF
Jiaxin Wen, Ruiqi Zhong, Akbir Khan, Ethan Perez, Jacob Steinhardt, Minlie Huang, Samuel R. Bowman, He He, and Shi Feng. 2025 · 2025
Closest in time.
Livebench: A challenging, contamination-free LLM benchmark
Colin White, Samuel Dooley, Manley Roberts, Arka Pal, Benjamin Feuer, Siddhartha Jain, Ravid Shwartz-Ziv, Neel Jain, Khalid Saifullah, Sreemanti Dey, Shubh-Agrawal, Sandeep Singh Sandha, Siddartha Venkat Naidu, Chinmay Hegde, Yann LeCun, Tom Goldstein, Willie Neiswanger, and Micah Goldblum. 2025 · 2025
Closest in time.
Cheating automatic LLM benchmarks: Null models achieve high win rates
Xiaosen Zheng, Tianyu Pang, Chao Du, Qian Liu, Jing Jiang, and Min Lin. 2025 · 2025
Closest in time.