Fetching the paper…
Reading the bibliography…
As large language models (LLMs) have grown in prevalence, particular benchmarks have become essential for the evaluation of these models and for understanding model capabilities.
Correlation calculated from faulty data
Charles Spearman. 1910 · 1910
Earlier work this paper cites.
Comparable tests and reliability
Jack W Dunlap. 1933 · 1933
Earlier work this paper cites.
Statistical methods for assessing agreement between two methods of clinical measurement
J Martin Bland and DouglasG Altman. 1986 · 1986
Earlier work this paper cites.
The handbook of questionnaire design
JA Krosnick and LR Fabrigar. 1991 · 1991
Earlier work this paper cites.
Numerical answer options: Logical or random order?
Renee M Huntley and Catherine J Welch. 1993 · 1993
Earlier work this paper cites.
A review of multiple-choice item-writing guidelines for classroom assessment
Thomas M Haladyna, Steven M Downing, and Michael C Rodriguez. 2002 · 2002
Earlier work this paper cites.
Does the answer order matter on multiple-choice exams?
Joel Tellinghuisen and Michelle M Sulikowski. 2008 · 2008
Earlier work this paper cites.
Large language models sensitivity to the order of options in multiple-choice questions
Pouya Pezeshkpour and Estevam Hruschka. 2024 · 2017
Earlier work this paper cites.
Synthetic and natural noise both break neural machine translation
Yonatan Belinkov and Yonatan Bisk. 2018 · 2018
Earlier work this paper cites.
HotFlip: White-box adversarial examples for text classification
Javid Ebrahimi, Anyi Rao, Daniel Lowd, and Dejing Dou. 2018 · 2018
Earlier work this paper cites.
Improving the robustness of question answering systems to question paraphrasing
Wee Chung Gan and Hwee Tou Ng. 2019 · 2019
Earlier work this paper cites.
Linking artificial and human neural representations of language
Jon Gauthier and Roger Levy. 2019 · 2019
Earlier work this paper cites.
How can we know what language models know?
Zhengbao Jiang, Frank F. Xu, Jun Araki, and Graham Neubig. 2020 · 2020
Earlier work this paper cites.
Beyond accuracy: Behavioral testing of NLP models with CheckList
Marco Tulio Ribeiro, Tongshuang Wu, Carlos Guestrin, and Sameer Singh. 2020 · 2020
Earlier work this paper cites.
What do we expect from multiple-choice QA systems?
Krunal Shah, Nitish Gupta, and Dan Roth. 2020 · 2020
Earlier work this paper cites.
Assessing the benchmarking capacity of machine reading comprehension datasets
Saku Sugawara, Pontus Stenetorp, Kentaro Inui, and Akiko Aizawa. 2020 · 2020
Earlier work this paper cites.
Making pre-trained language models better few-shot learners
Tianyu Gao, Adam Fisch, and Danqi Chen. 2021 · 2021
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021 · 2021
Earlier work this paper cites.
Contextualized perturbation for textual adversarial attack
Dianqi Li, Yizhe Zhang, Hao Peng, Liqun Chen, Chris Brockett, Ming-Ting Sun, and Bill Dolan. 2021 · 2021
Cited alongside, same era.
Evaluating the robustness of neural language models to input perturbations
Milad Moradi and Matthias Samwald. 2021 · 2021
Cited alongside, same era.
Masked language modeling and the distributional hypothesis: Order word matters pre-training for little
Koustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau, Adina Williams, and Douwe Kiela. 2021a · 2021
Cited alongside, same era.
Does the response options placement provide clues to the correct answers in multiple-choice tests? a systematic review
Séverin Lions, Carlos Monsalve, Pablo Dartnell, María Paz Blanco, Gabriel Ortega, and Julie Lemarié. 2022 · 2022
Cited alongside, same era.
Augly: Data augmentations for robustness
Zoe Papakipos and Joanna Bitton. 2022 · 2022
Cited alongside, same era.
Artifacts or abduction: How do LLMs answer multiple-choice questions without the question?
Nishant Balepur, Abhilasha Ravichander, and Rachel Rudinger. 2024 · 2024
Closest in time.
Is your large language model knowledgeable or a choices-only cheater?
Nishant Balepur and Rachel Rudinger. 2024 · 2024
Closest in time.
Aryo Pradipta Gema, Joshua Ong Jun Leang, Giwon Hong, Alessio Devoto, Alberto Carlo Maria Mancino, Rohit Saxena, Xuanli He, Yu Zhao, Xiaotang Du, Mohammad Reza Ghasemi Madani, et al. 2024 · 2024
Closest in time.
Reverse training to nurse the reversal curse
Olga Golovneva, Zeyuan Allen-Zhu, Jason Weston, and Sainbayar Sukhbaatar. 2024 · 2024
Closest in time.
Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, Arthur Mensch, Blanche Savary, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Emma Bou Hanna, Florian Bressand, Gianna Lengyel, Guillaume Bour, Guillaume Lample, Lélio Renard Lavaud, Lucile Saulnier, Marie-Anne Lachaux, Pierre Stock, Sandeep Subramanian, Sophia Yang, Szymon Antoniak, Teven Le Scao, Théophile Gervet, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2024 · 2024
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Perturbation augmentation for fairer nlp
Rebecca Qian, Candace Ross, Jude Fernandes, Eric Smith, Douwe Kiela, and Adina Williams. 2022 · 2022
Cited alongside, same era.
The curious case of absolute position embeddings
Koustuv Sinha, Amirhossein Kazemnejad, Siva Reddy, Joelle Pineau, Dieuwke Hupkes, and Adina Williams. 2022 · 2022
Cited alongside, same era.
The falcon series of open language models
Ebtesam Almazrouei, Hamza Alobeidli, Abdulaziz Alshamsi, Alessandro Cappelli, Ruxandra Cojocaru, Mérouane Debbah, Étienne Goffinet, Daniel Hesslow, Julien Launay, Quentin Malartic, Daniele Mazzotta, Badreddine Noune, Baptiste Pannier, and Guilherme Penedo. 2023 · 2023
Cited alongside, same era.
The reversal curse: Llms trained on" a is b" fail to learn" b is a"
Lukas Berglund, Meg Tong, Max Kaufmann, Mikita Balesni, Asa Cooper Stickland, Tomasz Korbak, and Owain Evans. 2023 · 2023
Cited alongside, same era.
A framework for few-shot language model evaluation
Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. 2023 · 2023
Cited alongside, same era.
Robustness of named-entity replacements for in-context learning
Saeed Goodarzi, Nikhil Kagita, Dennis Minn, Shufan Wang, Roberto Dessi, Shubham Toshniwal, Adina Williams, Jack Lanchantin, and Koustuv Sinha. 2023 · 2023
Cited alongside, same era.
Albert Q. Jiang, Alexandre Sablayrolles, Arthur Mensch, Chris Bamford, Devendra Singh Chaplot, Diego de las Casas, Florian Bressand, Gianna Lengyel, Guillaume Lample, Lucile Saulnier, Lélio Renard Lavaud, Marie-Anne Lachaux, Pierre Stock, Teven Le Scao, Thibaut Lavril, Thomas Wang, Timothée Lacroix, and William El Sayed. 2023 · 2023
Cited alongside, same era.
Closest in time.
The factorization curse: Which tokens you predict underlie the reversal curse and more
Ouail Kitouni, Niklas Nolte, Diane Bouchacourt, Adina Williams, Mike Rabbat, and Mark Ibrahim. 2024 · 2024
Closest in time.
Anchored answers: Unravelling positional bias in gpt-2’s multiple-choice questions
Ruizhe Li and Yanjun Gao. 2024 · 2024
Closest in time.
Set-based prompting: Provably solving the language model order dependency problem
Reid McIlroy-Young, Katrina Brown, Conlan Olson, Linjun Zhang, and Cynthia Dwork. 2024 · 2024
Closest in time.
Introducing meta llama 3: The most capable openly available llm to date
Meta. 2024 · 2024
Closest in time.
Beyond performance: Quantifying and mitigating label bias in LLMs
Yuval Reif and Roy Schwartz. 2024 · 2024
Closest in time.
What makes a good metric? evaluating automatic metrics for text-to-image consistency
Candace Ross, Melissa Hall, Adriana Romero-Soriano, and Adina Williams. 2024 · 2024
Closest in time.
Paul Röttger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck, Hannah Rose Kirk, Hinrich Schütze, and Dirk Hovy. 2024 · 2024
Closest in time.
Gemma: Open models based on gemini research and technology
Gemma Team, Thomas Mesnard, Cassidy Hardin, Robert Dadashi, Surya Bhupatiraju, Shreya Pathak, Laurent Sifre, Morgane Rivière, Mihir Sanjay Kale, Juliette Love, Pouya Tafti, Léonard Hussenot, Pier Giuseppe Sessa, Aakanksha Chowdhery, Adam Roberts, Aditya Barua, Alex Botev, Alex Castro-Ros, Ambrose Slone, Amélie Héliou, Andrea Tacchetti, Anna Bulanova, Antonia Paterson, Beth Tsai, Bobak Shahriari, Charline Le Lan, Christopher A. Choquette-Choo, Clément Crepy, Daniel Cer, Daphne Ippolito, David Reid, Elena Buchatskaya, Eric Ni, Eric Noland, Geng Yan, George Tucker, George-Christian Muraru, Grigory Rozhdestvenskiy, Henryk Michalewski, Ian Tenney, Ivan Grishchenko, Jacob Austin, James Keeling, Jane Labanowski, Jean-Baptiste Lespiau, Jeff Stanway, Jenny Brennan, Jeremy Chen, Johan Ferret, Justin Chiu, Justin Mao-Jones, Katherine Lee, Kathy Yu, Katie Millican, Lars Lowe Sjoesund, Lisa Lee, Lucas Dixon, Machel Reid, Maciej Mikuła, Mateo Wirth, Michael Sharman, Nikolai Chinaev, Nithum Thain, Olivier Bachem, Oscar Chang, Oscar Wahltinez, Paige Bailey, Paul Michel, Petko Yotov, Rahma Chaabouni, Ramona Comanescu, Reena Jana, Rohan Anil, Ross McIlroy, Ruibo Liu, Ryan Mullins, Samuel L Smith, Sebastian Borgeaud, Sertan Girgin, Sholto Douglas, Shree Pandya, Siamak Shakeri, Soham De, Ted Klimenko, Tom Hennigan, Vlad Feinberg, Wojciech Stokowiec, Yu hui Chen, Zafarali Ahmed, Zhitao Gong, Tris Warkentin, Ludovic Peran, Minh Giang, Clément Farabet, Oriol Vinyals, Jeff Dean, Koray Kavukcuoglu, Demis Hassabis, Zoubin Ghahramani, Douglas Eck, Joelle Barral, Fernando Pereira, Eli Collins, Armand Joulin, Noah Fiedel, Evan Senter, Alek Andreev, and Kathleen Kenealy. 2024 · 2024
Closest in time.
Unveiling selection biases: Exploring order and token sensitivity in large language models
Sheng-Lun Wei, Cheng-Kuang Wu, Hen-Hsen Huang, and Hsin-Hsi Chen. 2024 · 2024
Closest in time.
Llms’ classification performance is overclaimed
Hanzi Xu, Renze Lou, Jiangshu Du, Vahid Mahzoon, Elmira Talebianaraki, Zhuoan Zhou, Elizabeth Garrison, Slobodan Vucetic, and Wenpeng Yin. 2024 · 2024
Closest in time.
Strengthened symbol binding makes large language models reliable multiple-choice selectors
Mengge Xue, Zhenyu Hu, Meng Zhao, Liqun Liu, Kuo Liao, Shuang Li, Honglin Han, and Chengguo Yin. 2024 · 2024
Closest in time.
Revisiting the self-consistency challenges in multi-choice question formats for large language model evaluation
Wenjie Zhou, Qiang Wang, Mingzhou Xu, Ming Chen, and Xiangyu Duan. 2024 · 2024
Closest in time.
Fool your large (vision and) language models with embarrassingly simple permutations
Yongshuo Zong, Yist Tingyang YU, Bingchen Zhao, Ruchika Chavhan, and Timothy Hospedales. 2024 · 2024
Closest in time.