Fetching the paper…
Reading the bibliography…
The increasing versatility of language models (LMs) has given rise to a new class of benchmarks that comprehensively assess a broad range of capabilities.
The theory of the estimation of test reliability
G Frederic Kuder and Marion W Richardson. 1937 · 1937
Earlier work this paper cites.
Response sets and test validity
Lee J Cronbach. 1946 · 1946
Earlier work this paper cites.
A difficulty in the concept of social welfare
Kenneth J Arrow. 1950 · 1950
Earlier work this paper cites.
Beyond user self-reported likert scale ratings: A comparison model for automatic dialog evaluation
Weixin Liang, James Zou, and Zhou Yu. 2020 · 2005
Earlier work this paper cites.
Measuring massive multitask language understanding
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Xiaodong Song, and Jacob Steinhardt. 2020 · 2009
Earlier work this paper cites.
The original borda count and partial voting
Peter Emerson. 2013 · 2013
Earlier work this paper cites.
A weighted correlation index for rankings with ties
Sebastiano Vigna. 2014 · 2014
Earlier work this paper cites.
How to build a benchmark
Jóakim v. Kistowski, Jeremy A Arnold, Karl Huppler, Klaus-Dieter Lange, John L Henning, and Paul Cao. 2015 · 2015
Earlier work this paper cites.
Findings of the 2016 conference on machine translation
Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Aurélie Névéol, Mariana Neves, Martin Popel, Matt Post, Raphael Rubino, Carolina Scarton, Lucia Specia, Marco Turchi, Karin Verspoor, and Marcos Zampieri. 2016 · 2016
Earlier work this paper cites.
Sequence effects in crowdsourced annotations
Nitika Mathur, Timothy Baldwin, and Trevor Cohn. 2017 · 2017
Earlier work this paper cites.
Annotation artifacts in natural language inference data
Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A. Smith. 2018 · 2018
Earlier work this paper cites.
Can a suit of armor conduct electricity? a new dataset for open book question answering
Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018 · 2018
Earlier work this paper cites.
Hypothesis only baselines in natural language inference
Adam Poliak, Jason Naradowsky, Aparajita Haldar, Rachel Rudinger, and Benjamin Van Durme. 2018 · 2018
Earlier work this paper cites.
BLEU is not suitable for the evaluation of text simplification
Elior Sulem, Omri Abend, and Ari Rappoport. 2018 · 2018
Earlier work this paper cites.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 · 2018
Earlier work this paper cites.
Low-bit quantization of neural networks for efficient inference
Yoni Choukroun, Eli Kravchik, Fan Yang, and Pavel Kisilev. 2019 · 2019
Earlier work this paper cites.
Summeval: Re-evaluating summarization evaluation
A. R. Fabbri, Wojciech Kryscinski, Bryan McCann, Richard Socher, and Dragomir R. Radev. 2020 · 2020
Earlier work this paper cites.
Tangled up in BLEU: Reevaluating the evaluation of automatic machine translation evaluation metrics
Nitika Mathur, Timothy Baldwin, and Trevor Cohn. 2020 · 2020
Cited alongside, same era.
oLMpics-on what language model pre-training captures
Alon Talmor, Yanai Elazar, Yoav Goldberg, and Jonathan Berant. 2020 · 2020
Cited alongside, same era.
RAFT: A real-world few-shot text classification benchmark
Neel Alex, Eli Lifland, Lewis Tunstall, Abhishek Thakur, Pegah Maham, C. Jess Riedel, Emmie Hine, Carolyn Ashurst, Paul Sedille, Alexis Carlier, Michael Noetel, and Andreas Stuhlmüller. 2021 · 2021
Cited alongside, same era.
The devil is in the detail: Simple tricks improve systematic generalization of transformers
Róbert Csordás, Kazuki Irie, and Juergen Schmidhuber. 2021 · 2021
Cited alongside, same era.
A framework for few-shot language model evaluation
Leo Gao, Jonathan Tow, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Kyle McDonell, Niklas Muennighoff, Jason Phang, Laria Reynolds, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. 2021 · 2021
Cited alongside, same era.
Training language models to follow instructions with human feedback
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. 2022 · 2022
Later among the works it cites.
CometKiwi: IST-unbabel 2022 submission for the quality estimation shared task
Ricardo Rei, Marcos Treviso, Nuno M. Guerreiro, Chrysoula Zerva, Ana C Farinha, Christine Maroti, José G. C. de Souza, Taisiya Glushkova, Duarte Alves, Luisa Coheur, Alon Lavie, and André F. T. Martins. 2022 · 2022
Later among the works it cites.
Findings of the WMT 2022 shared task on quality estimation
Chrysoula Zerva, Frédéric Blain, Ricardo Rei, Piyawat Lertvittayakumjorn, José G. C. de Souza, Steffen Eger, Diptesh Kanojia, Duarte Alves, Constantin Orăsan, Marina Fomicheva, André F. T. Martins, and Lucia Specia. 2022 · 2022
Later among the works it cites.
Beyond the imitation game: Quantifying and extrapolating the capabilities of language models
BIG bench authors. 2023 · 2023
Closest in time.
Emergent and predictable memorization in large language models
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The GEM benchmark: Natural language generation, its evaluation and metrics
Sebastian Gehrmann, Tosin Adewumi, Karmanya Aggarwal, Pawan Sasanka Ammanamanchi, Anuoluwapo Aremu, Antoine Bosselut, Khyathi Raghavi Chandu, Miruna-Adriana Clinciu, Dipanjan Das, Kaustubh Dhole, et al. 2021 · 2021
Cited alongside, same era.
q 2 q^{2} : Evaluating factual consistency in knowledge-grounded dialogues via question generation and question answering
Or Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman, Idan Szpektor, and Omri Abend. 2021 · 2021
Cited alongside, same era.
On the stability of system rankings at WMT
Rebecca Knowles. 2021 · 2021
Cited alongside, same era.
To ship or not to ship: An extensive evaluation of automatic metrics for machine translation
Tom Kocmi, Christian Federmann, Roman Grundkiewicz, Marcin Junczys-Dowmunt, Hitokazu Matsushita, and Arul Menezes. 2021 · 2021
Cited alongside, same era.
Better than average: Paired evaluation of NLP systems
Maxime Peyrard, Wei Zhao, Steffen Eger, and Robert West. 2021 · 2021
Cited alongside, same era.
Where to start? analyzing the potential value of intermediate models
Leshem Choshen, Elad Venezian, Shachar Don-Yehia, Noam Slonim, and Yoav Katz. 2022 · 2022
Cited alongside, same era.
Results of WMT22 metrics shared task: Stop using BLEU – neural metrics are better and more robust
Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, Eleftherios Avramidis, Tom Kocmi, George Foster, Alon Lavie, and André F. T. Martins. 2022 · 2022
Cited alongside, same era.
Stella Biderman, USVSN Sai Prashanth, Lintang Sutawika, Hailey Schoelkopf, Quentin Anthony, Shivanshu Purohit, and Edward Raf. 2023 · 2023
Closest in time.
A survey on evaluation of large language models
Yupeng Chang, Xu Wang, Jindong Wang, Yuan Wu, Kaijie Zhu, Hao Chen, Linyi Yang, Xiaoyuan Yi, Cunxiang Wang, Yidong Wang, et al. 2023 · 2023
Closest in time.
Accelerating large language model decoding with speculative sampling
Charlie Chen, Sebastian Borgeaud, Geoffrey Irving, Jean-Baptiste Lespiau, Laurent Sifre, and John Jumper. 2023 · 2023
Closest in time.
Instructeval: Towards holistic evaluation of instruction-tuned large language models
Yew Ken Chia, Pengfei Hong, Lidong Bing, and Soujanya Poria. 2023 · 2023
Closest in time.
Opencompass: A universal evaluation platform for foundation models
OpenCompass Contributors. 2023 · 2023
Closest in time.
Why can gpt learn in-context? language models implicitly perform gradient descent as meta-optimizers
Damai Dai, Yutao Sun, Li Dong, Yaru Hao, Shuming Ma, Zhifang Sui, and Furu Wei. 2023 · 2023
Closest in time.
Scaling down to scale up: A guide to parameter-efficient fine-tuning
Vladislav Lialin, Vijeta Deshpande, and Anna Rumshisky. 2023 · 2023
Closest in time.
Benchmarking large language model capabilities for conditional generation
Joshua Maynez, Priyanka Agrawal, and Sebastian Gehrmann. 2023 · 2023
Closest in time.
What In-Context Learning “Learns” In-Context: Disentangling Task Recognition and Task Learning
Jane Pan. 2023 · 2023
Closest in time.
Evaluating instruction-tuned large language models on code comprehension and generation
Zhiqiang Yuan, Junwei Liu, Qiancheng Zi, Mingwei Liu, Xin Peng, and Yiling Lou. 2023 · 2023
Closest in time.
DialogStudio: Towards richest and most diverse unified dataset collection for conversational AI
Jianguo Zhang, Kun Qian, Zhiwei Liu, Shelby Heinecke, Rui Meng, Ye Liu, Zhou Yu, Silvio Savarese, and Caiming Xiong. 2023 · 2023
Closest in time.
Unitxt: Flexible, shareable and reusable data preparation and evaluation for generative ai
Elron Bandel, Yotam Perlitz, Elad Venezian, Roni Friedman-Melamed, Ofir Arviv, Matan Orbach, Shachar Don-Yehiya, Dafna Sheinwald, Ariel Gera, Leshem Choshen, Michal Shmueli-Scheuer, and Yoav Katz. 2024 · 2024
Closest in time.