Fetching the paper…
Reading the bibliography…
Evaluation in NLP is usually done by comparing the scores of competing systems independently averaged over a common set of test instances.
A law of comparative judgement
Louis Leon Thurstone. 1927 · 1927
Earlier work this paper cites.
Better summarization evaluation with word embeddings for ROUGE
Jun-Ping Ng and Viktoria Abrecht. 2015 · 1930
Earlier work this paper cites.
The Design of Experiments
Ronald A. Fisher. 1935 · 1935
Earlier work this paper cites.
Individual comparisons by ranking methods
Frank Wilcoxon. 1945 · 1945
Earlier work this paper cites.
A difficulty in the concept of social welfare
Kenneth J. Arrow. 1950 · 1950
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E. Terry. 1952 · 1952
Earlier work this paper cites.
The rating of chessplayers, past and present
Arpad E. Elo. 1978 · 1978
Earlier work this paper cites.
BLEU: A Method for Automatic Evaluation of Machine Translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
An information-theoretic approach to automatic evaluation of summaries
Chin-Yew Lin, Guihong Cao, Jianfeng Gao, and Jian-Yun Nie. 2006 · 2006
Earlier work this paper cites.
Trueskill™: A bayesian skill rating system
Ralf Herbrich, Tom Minka, and Thore Graepel. 2007 · 2007
Earlier work this paper cites.
Meteor: An Automatic Metric for MT Evaluation with High Levels of Correlation with Human Judgments
Alon Lavie and Abhaya Agarwal. 2007 · 2007
Earlier work this paper cites.
Ranking human and machine summarization systems
Peter Rankel, John Conroy, Eric Slud, and Dianne O’Leary. 2011 · 2011
Earlier work this paper cites.
An Assessment of the Accuracy of Automatic Evaluation in Summarization
Karolina Owczarzak, John M. Conroy, Hoa Trang Dang, and Ani Nenkova. 2012 · 2012
Cited alongside, same era.
Microsoft coco: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014 · 2014
Cited alongside, same era.
Efficient elicitation of annotations for human evaluation of machine translation
Keisuke Sakaguchi, Matt Post, and Benjamin Van Durme. 2014 · 2014
Cited alongside, same era.
Squibs: Evaluating human pairwise preference judgments
Mark Dras. 2015 · 2015
Cited alongside, same era.
Re-evaluating Automatic Summarization with BLEU and 192 Shades of ROUGE
Yvette Graham. 2015 · 2015
Cited alongside, same era.
Teaching machines to read and comprehend
Karl Moritz Hermann, Tomas Kocisky, Edward Grefenstette, Lasse Espeholt, Will Kay, Mustafa Suleyman, and Phil Blunsom. 2015 · 2015
Results of the WMT18 metrics shared task: Both characters and embeddings achieve good performance
Qingsong Ma, Ondřej Bojar, and Yvette Graham. 2018 · 2018
Later among the works it cites.
RankME: Reliable human ratings for natural language generation
Jekaterina Novikova, Ondřej Dušek, and Verena Rieser. 2018 · 2018
Later among the works it cites.
Learning latent parameters without human response patterns: Item response theory with artificial crowds
John P. Lalor, Hao Wu, and Hong Yu. 2019 · 2019
Later among the works it cites.
Results of the WMT19 Metrics Shared Task: Segment-Level and Strong MT Systems Pose Big Challenges
Qingsong Ma, Johnny Wei, Ondřej Bojar, and Yvette Graham. 2019 · 2019
Later among the works it cites.
MoverScore: Text generation evaluating with contextualized embeddings and earth mover distance
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
CIDEr: Consensus-based Image Description Evaluation
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Cited alongside, same era.
Building an evaluation scale using item response theory
John P. Lalor, Hao Wu, and Hong Yu. 2016 · 2016
Cited alongside, same era.
Results of the WMT17 metrics shared task
Ondřej Bojar, Yvette Graham, and Amir Kamran. 2017 · 2017
Cited alongside, same era.
Learning to score system summaries for better content selection evaluation
Maxime Peyrard, Teresa Botschen, and Iryna Gurevych. 2017 · 2017
Cited alongside, same era.
chrF++: Words Helping Character n-grams
Maja Popovic. 2017 · 2017
Cited alongside, same era.
The wilcoxon–mann–whitney procedure fails as a test of medians
George W. Divine, H. James Norton, Anna E. Barón, and Elizabeth Juarez-Colunga. 2018 · 2018
Cited alongside, same era.
Spot the bot: A robust and efficient framework for the evaluation of conversational dialogue systems
Jan Deriu, Don Tuggener, Pius von Däniken, Jon Ander Campos, Alvaro Rodrigo, Thiziri Belkacem, Aitor Soroa, Eneko Agirre, and Mark Cieliebak. 2020 · 2020
Later among the works it cites.
BLEU might be guilty but references are not innocent
Markus Freitag, David Grangier, and Isaac Caswell. 2020 · 2020
Later among the works it cites.
Statistical power and translationese in machine translation evaluation
Yvette Graham, Barry Haddow, and Philipp Koehn. 2020 · 2020
Later among the works it cites.
Tangled up in BLEU: Reevaluating the evaluation of automatic machine translation evaluation metrics
Nitika Mathur, Timothy Baldwin, and Trevor Cohn. 2020 · 2020
Later among the works it cites.
USR: An unsupervised and reference free evaluation metric for dialog generation
Shikib Mehri and Maxine Eskenazi. 2020 · 2020
Later among the works it cites.
Item response theory for efficient human evaluation of chatbots
João Sedoc and Lyle Ungar. 2020 · 2020
Later among the works it cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Later among the works it cites.
On the limitations of cross-lingual encoders as exposed by reference-free machine translation evaluation
Wei Zhao, Goran Glavaš, Maxime Peyrard, Yang Gao, Robert West, and Steffen Eger. 2020 · 2020
Later among the works it cites.