Fetching the paper…
Reading the bibliography…
In Machine Learning, a benchmark refers to an ensemble of datasets associated with one or multiple metrics together with a way to aggregate different systems performances.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi · 1904
Earlier work this paper cites.
Paws: Paraphrase adversaries from word scrambling
Yuan Zhang, Jason Baldridge, and Luheng He · 1904
Earlier work this paper cites.
A new measure of rank correlation
Maurice G Kendall · 1938
Earlier work this paper cites.
The treatment of ties in ranking problems
Maurice G Kendall · 1945
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry · 1952
Earlier work this paper cites.
Individual choice behavior: A theoretical analysis, 1959
Luce R Duncan · 1959
Earlier work this paper cites.
Mathematics without numbers
John G Kemeny · 1959
Earlier work this paper cites.
When Do Noisy Votes Reveal the Truth?
Ioannis Caragiannis, Ariel D Procaccia, and Nisarg Shah · 1962
Earlier work this paper cites.
The analysis of permutations
Robin L Plackett · 1975
Earlier work this paper cites.
A consistent extension of condorcet’s election principle
H Peyton Young and Arthur Levenglick · 1978
Earlier work this paper cites.
Distance based ranking models
Michael A Fligner and Joseph S Verducci · 1986
Earlier work this paper cites.
Multistage Ranking Models
Michael A Fligner and Joseph S Verducci · 1988
Earlier work this paper cites.
Condorcet’s theory of voting
H Peyton Young · 1988
Earlier work this paper cites.
The computational difficulty of manipulating an election
John J Bartholdi, Craig A Tovey, and Michael A Trick · 1989
Earlier work this paper cites.
Fundamentals of social choice theory
Roger B Myerson · 1996
Earlier work this paper cites.
Rank aggregation methods for the web
Cynthia Dwork, Ravi Kumar, Moni Naor, and Dandapani Sivakumar · 2001
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu · 2002
Earlier work this paper cites.
Comparing top k lists
Ronald Fagin, Ravi Kumar, and Dakshinamurthi Sivakumar · 2003
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie · 2005
Earlier work this paper cites.
The pascal recognising textual entailment challenge
Ido Dagan, Oren Glickman, and Bernardo Magnini · 2005
Earlier work this paper cites.
Automatically constructing a corpus of sentential paraphrases
William B Dolan and Chris Brockett · 2005
Earlier work this paper cites.
Ordering by weighted number of wins gives a good ranking for weighted tournaments
Don Coppersmith, Lisa Fleischer, and Atri Rudra · 2006
Earlier work this paper cites.
An information-theoretic approach to automatic evaluation of summaries
Chin-Yew Lin, Guihong Cao, Jianfeng Gao, and Jian-Yun Nie · 2006
Earlier work this paper cites.
The third pascal recognizing textual entailment challenge
Danilo Giampiccolo, Bernardo Magnini, Ido Dagan, and William B Dolan · 2007
Earlier work this paper cites.
Overview of the tac 2008 update summarization task
Hoa Trang Dang, Karolina Owczarzak, et al · 2008
Earlier work this paper cites.
Empirical bernstein stopping
Volodymyr Mnih, Csaba Szepesvári, and Jean-Yves Audibert · 2008
Earlier work this paper cites.
The fifth pascal recognizing textual entailment challenge
Luisa Bentivogli, Peter Clark, Ido Dagan, and Danilo Giampiccolo · 2009
Earlier work this paper cites.
How similarity helps to efficiently compute kemeny rankings
Nadja Betzler, Michael R Fellows, Jiong Guo, Rolf Niedermeier, and Frances A Rosamond · 2009
Earlier work this paper cites.
Bypassing combinatorial protections: Polynomial-time algorithms for single-peaked electorates
Felix Brandt, Markus Brill, Edith Hemaspaandra, and Lane A Hemaspaandra · 2010
Earlier work this paper cites.
Overview of the tac 2011 summarization track: Guided task and aesop task
Karolina Owczarzak and Hoa Trang Dang · 2011
Earlier work this paper cites.
Choice of plausible alternatives: An evaluation of commonsense causal reasoning
Melissa Roemmele, Cosmin Adrian Bejan, and Andrew S Gordon · 2011
Earlier work this paper cites.
Long range arena: A benchmark for efficient transformers
Yi Tay, Mostafa Dehghani, Samira Abnar, Yikang Shen, Dara Bahri, Philip Pham, Jinfeng Rao, Liu Yang, Sebastian Ruder, and Donald Metzler · 2011
Earlier work this paper cites.
Experiments with kemeny ranking: What works when?
Alnur Ali and Marina Meilă · 2012
Cited alongside, same era.
The winograd schema challenge
Hector Levesque, Ernest Davis, and Leora Morgenstern · 2012
Cited alongside, same era.
Recursive deep models for semantic compositionality over a sentiment treebank
Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D Manning, Andrew Y Ng, and Christopher Potts · 2013
Cited alongside, same era.
Meteor universal: Language specific translation evaluation for any target language
Michael Denkowski and Alon Lavie · 2014
Cited alongside, same era.
From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Peter Young, Alice Lai, Micah Hodosh, and Julia Hockenmaier · 2014
Cited alongside, same era.
Better summarization evaluation with word embeddings for rouge
Mlqa: Evaluating cross-lingual extractive question answering
Patrick Lewis, Barlas Oğuz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk · 2019
Later among the works it cites.
Massively multilingual transfer for ner
Afshin Rahimi, Yuan Li, and Trevor Cohn · 2019
Later among the works it cites.
Superglue: A stickier benchmark for general-purpose language understanding systems
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R Bowman · 2019
Later among the works it cites.
Neural network acceptability judgments
Alex Warstadt, Amanpreet Singh, and Samuel R Bowman · 2019
Later among the works it cites.
Paws-x: A cross-lingual adversarial dataset for paraphrase identification
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jun-Ping Ng and Viktoria Abrecht · 2015
Cited alongside, same era.
chrf: character n-gram f-score for automatic mt evaluation
Maja Popović · 2015
Cited alongside, same era.
Squad: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang · 2016
Cited alongside, same era.
Mastering the game of go with deep neural networks and tree search
David Silver, Aja Huang, Chris J Maddison, Arthur Guez, Laurent Sifre, George Van Den Driessche, Julian Schrittwieser, Ioannis Antonoglou, Veda Panneershelvam, Marc Lanctot, et al · 2016
Cited alongside, same era.
Semeval-2017 task 1: Semantic textual similarity-multilingual and cross-lingual focused evaluation
Daniel Cer, Mona Diab, Eneko Agirre, Inigo Lopez-Gazpio, and Lucia Specia · 2017
Cited alongside, same era.
Creating training corpora for nlg micro-planning
Claire Gardent, Anastasia Shimorina, Shashi Narayan, and Laura Perez-Beltrachini · 2017
Cited alongside, same era.
Generalization in deep learning
Kenji Kawaguchi, Leslie Pack Kaelbling, and Yoshua Bengio · 2017
Cited alongside, same era.
Yinfei Yang, Yuan Zhang, Chris Tar, and Jason Baldridge · 2019
Later among the works it cites.
Adversarial examples: Attacks and defenses for deep learning
Xiaoyong Yuan, Pan He, Qile Zhu, and Xiaolin Li · 2019
Later among the works it cites.
Moverscore: Text generation evaluating with contextualized embeddings and earth mover distance
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M Meyer, and Steffen Eger · 2019
Later among the works it cites.
Re-evaluating evaluation in text summarization
Manik Bhandari, Pranav Gour, Atabak Ashfaq, Pengfei Liu, and Graham Neubig · 2020
Later among the works it cites.
Hierarchical pre-training for sequence labelling in spoken dialog
Emile Chapuis, Pierre Colombo, Matteo Manica, Matthieu Labeau, and Chloe Clavel · 2020
Later among the works it cites.
Tydi qa: A benchmark for information-seeking question answering in typologically diverse languages
Jonathan H Clark, Eunsol Choi, Michael Collins, Dan Garrette, Tom Kwiatkowski, Vitaly Nikolaev, and Jennimaria Palomaki · 2020
Later among the works it cites.
Guiding attention in sequence-to-sequence models for dialogue act prediction
Pierre Colombo, Emile Chapuis, Matteo Manica, Emmanuel Vignon, Giovanna Varni, and Chloe Clavel · 2020
Later among the works it cites.
The importance of fillers for text representations of speech transcripts
Tanvi Dinkar, Pierre Colombo, Matthieu Labeau, and Chloé Clavel · 2020
Later among the works it cites.
Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson · 2020
Later among the works it cites.
Usr: An unsupervised and reference free evaluation metric for dialog generation
Shikib Mehri and Maxine Eskenazi · 2020
Later among the works it cites.
Xtreme: A massively multilingual multi-task benchmark for evaluating cross-lingual generalization
Aditya Siddhant, Junjie Hu, Melvin Johnson, Orhan Firat, and Sebastian Ruder · 2020
Later among the works it cites.
Ext5: Towards extreme multi-task scaling for transfer learning
Vamsi Aribandi, Yi Tay, Tal Schuster, Jinfeng Rao, Huaixiu Steven Zheng, Sanket Vaibhav Mehta, Honglei Zhuang, Vinh Q Tran, Dara Bahri, Jianmo Ni, et al · 2021
Later among the works it cites.
Private and non-private uniformity testing for ranking data
Róbert Busa-Fekete, Dimitris Fotakis, and Emmanouil Zampetakis · 2021
Later among the works it cites.
Code-switched inspired losses for generic spoken dialog representations
Emile Chapuis, Pierre Colombo, Matthieu Labeau, and Chloe Clavel · 2021
Later among the works it cites.
Learning to represent and generate text using information measures
Pierre Colombo · 2021
Later among the works it cites.
Mostafa Dehghani, Yi Tay, Alexey A Gritsenko, Zhe Zhao, Neil Houlsby, Fernando Diaz, Donald Metzler, and Oriol Vinyals · 2021
Later among the works it cites.
Nl-augmenter: A framework for task-sensitive natural language augmentation
Kaustubh D Dhole, Varun Gangal, Sebastian Gehrmann, Aadesh Gupta, Zhenhao Li, Saad Mahamood, Abinaya Mahendiran, Simon Mille, Ashish Srivastava, Samson Tan, et al · 2021
Later among the works it cites.
Summeval: Re-evaluating summarization evaluation
Alexander R Fabbri, Wojciech Kryściński, Bryan McCann, Caiming Xiong, Richard Socher, and Dragomir Radev · 2021
Later among the works it cites.
Explainaboard: An explainable leaderboard for nlp
Pengfei Liu, Jinlan Fu, Yang Xiao, Weizhe Yuan, Shuaicheng Chang, Junqi Dai, Yixin Liu, Zihuiwen Ye, Zi-Yi Dou, and Graham Neubig · 2021
Later among the works it cites.
Reusable templates and guides for documenting datasets and models for natural language processing and generation: A case study of the HuggingFace and GEM data and model cards
Angelina McMillan-Major, Salomey Osei, Juan Diego Rodriguez, Pawan Sasanka Ammanamanchi, Sebastian Gehrmann, and Yacine Jernite · 2021
Later among the works it cites.
Better than average: Paired evaluation of nlp systems
Maxime Peyrard, Wei Zhao, Steffen Eger, and Robert West · 2021
Later among the works it cites.
Tharindu Ranasinghe, Constantin Orasan, and Ruslan Mitkov · 2021
Later among the works it cites.
Challenges and Opportunities in NLP Benchmarking
Sebastian Ruder · 2021
Later among the works it cites.
Multitask prompted training enables zero-shot task generalization
Victor Sanh, Albert Webson, Colin Raffel, Stephen H Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Teven Le Scao, Arun Raja, et al · 2021
Later among the works it cites.
Depth-based pseudo-metrics between probability distributions
Guillaume Staerman, Pavlo Mozharovskyi, Stéphan Clémençon, and Florence d’Alché Buc · 2021
Later among the works it cites.
Finetuned language models are zero-shot learners
Jason Wei, Maarten Bosma, Vincent Y Zhao, Kelvin Guu, Adams Wei Yu, Brian Lester, Nan Du, Andrew M Dai, and Quoc V Le · 2021
Later among the works it cites.
Bartscore: Evaluating generated text as text generation
Weizhe Yuan, Graham Neubig, and Pengfei Liu · 2021
Later among the works it cites.
Fewnlu: Benchmarking state-of-the-art methods for few-shot natural language understanding
Yanan Zheng, Jing Zhou, Yujie Qian, Ming Ding, Jian Li, Ruslan Salakhutdinov, Jie Tang, Sebastian Ruder, and Zhilin Yang · 2021
Later among the works it cites.
Are larger pretrained language models uniformly better? comparing performance at the instance level
Ruiqi Zhong, Dhruba Ghosh, Dan Klein, and Jacob Steinhardt · 2021
Later among the works it cites.