Fetching the paper…
Reading the bibliography…
Current practices in metric evaluation focus on one single dataset, e.g., Newstest dataset in each year's WMT Metrics Shared Task.
A new measure of rank correlation
Maurice G Kendall. 1938 · 1938
Earlier work this paper cites.
The kolmogorov-smirnov test for goodness of fit
Frank J Massey Jr. 1951 · 1951
Earlier work this paper cites.
Accelerated dp based search for statistical translation
Christoph Tillmann, Stephan Vogel, Hermann Ney, Arkaitz Zubiaga, and Hassan Sawaf. 1997 · 1997
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Statistical significance tests for machine translation evaluation
Philipp Koehn. 2004 · 2004
Earlier work this paper cites.
From wer and ril to mer and wil: improved evaluation measures for connected speech recognition
Andrew Cameron Morris, Viktoria Maier, and Phil Green. 2004 · 2004
Earlier work this paper cites.
Jing Pan, Vincent Pham, Mohan Dorairaj, Huigang Chen, and Jeong-Yoon Lee. 2020 · 2004
Earlier work this paper cites.
A study of translation edit rate with targeted human annotation
Matthew Snover, Bonnie Dorr, Rich Schwartz, Linnea Micciulla, and John Makhoul. 2006 · 2006
Earlier work this paper cites.
Randomized significance tests in machine translation
Yvette Graham, Nitika Mathur, and Timothy Baldwin. 2014 · 2014
Cited alongside, same era.
Fitting sentence level translation evaluation with many dense features
Miloš Stanojević and Khalil Sima’an. 2014 · 2014
Cited alongside, same era.
chrF: character n-gram F-score for automatic MT evaluation
Maja Popović. 2015 · 2015
Cited alongside, same era.
Ten years of wmt evaluation campaigns: Lessons learnt
Ondrej Bojar, Christian Federmann, Barry Haddow, Philipp Koehn, Matt Post, and Lucia Specia. 2016 · 2016
Cited alongside, same era.
Achieving accurate conclusions in evaluation of automatic machine translation metrics
Yvette Graham and Qun Liu. 2016 · 2016
Cited alongside, same era.
CharacTer: Translation edit rate on character level
Weiyue Wang, Jan-Thorsten Peter, Hendrik Rosendahl, and Hermann Ney. 2016 · 2016
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Results of the WMT19 metrics shared task: Segment-level and strong MT systems pose big challenges
Qingsong Ma, Johnny Wei, Ondřej Bojar, and Yvette Graham. 2019 · 2019
Later among the works it cites.
MoverScore: Text generation evaluating with contextualized embeddings and earth mover distance
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. 2019 · 2019
Later among the works it cites.
Transformers: State-of-the-art natural language processing
Thomas Wolf, Lysandre Debut, Victor Sanh, Julien Chaumond, Clement Delangue, Anthony Moi, Pierric Cistac, Tim Rault, Remi Louf, Morgan Funtowicz, Joe Davison, Sam Shleifer, Patrick von Platen, Clara Ma, Yacine Jernite, Julien Plu, Canwen Xu, Teven Le Scao, Sylvain Gugger, Mariama Drame, Quentin Lhoest, and Alexander Rush. 2020 · 2020
Later among the works it cites.
Bertscore: Evaluating text generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
chrF++: words helping character n-grams
Maja Popović. 2017 · 2017
Cited alongside, same era.
Results of the WMT18 metrics shared task: Both characters and embeddings achieve good performance
Qingsong Ma, Ondřej Bojar, and Yvette Graham. 2018 · 2018
Cited alongside, same era.
Results of the WMT13 metrics shared task
Matouš Macháček and Ondřej Bojar. 2013a
Cited in the paper.
Results of the WMT13 metrics shared task
Matouš Macháček and Ondřej Bojar. 2013b
Cited in the paper.
Tangled up in BLEU: Reevaluating the evaluation of automatic machine translation evaluation metrics
Nitika Mathur, Timothy Baldwin, and Trevor Cohn. 2020a
Cited in the paper.
Results of the wmt20 metrics shared task
Nitika Mathur, Johnny Wei, Markus Freitag, Qingsong Ma, and Ondřej Bojar. 2020b
Cited in the paper.
Later among the works it cites.
Results of the wmt21 metrics shared task: Evaluating metrics with expert-based human evaluations on ted and news domain
Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, George Foster, Alon Lavie, and Ondřej Bojar. 2021 · 2021
Later among the works it cites.
Scientific credibility of machine translation research: A meta-evaluation of 769 papers
Benjamin Marie, Atsushi Fujita, and Raphael Rubino. 2021 · 2021
Later among the works it cites.