Fetching the paper…
Reading the bibliography…
Automatic metrics are commonly used as the exclusive tool for declaring the superiority of one machine translation system's quality over another.
Individual comparisons of grouped data by ranking methods
Frank Wilcoxon. 1946 · 1946
Earlier work this paper cites.
On a test of whether one of two random variables is stochastically larger than the other
Henry B Mann and Donald R Whitney. 1947 · 1947
Earlier work this paper cites.
An introduction to the bootstrap
Bradley Efron and Robert J Tibshirani. 1994 · 1994
Earlier work this paper cites.
Bleu: a Method for Automatic Evaluation of Machine Translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
BERTScore: Evaluating Text Generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2002
Earlier work this paper cites.
Methods of meta-analysis: Correcting error and bias in research findings
John E Hunter and Frank L Schmidt. 2004 · 2004
Earlier work this paper cites.
Statistical significance tests for machine translation evaluation
Philipp Koehn. 2004 · 2004
Earlier work this paper cites.
On some pitfalls in automatic evaluation and significance testing for mt
Stefan Riezler and John T Maxwell III. 2005 · 2005
Earlier work this paper cites.
Re-evaluating the role of Bleu in machine translation research
Chris Callison-Burch, Miles Osborne, and Philipp Koehn. 2006 · 2006
Earlier work this paper cites.
A study of translation edit rate with targeted human annotation
Matthew Snover, Bonnie Dorr, Richard Schwartz, Linnea Micciulla, and John Makhoul. 2006 · 2006
Earlier work this paper cites.
(Meta-) evaluation of machine translation
Chris Callison-Burch, Cameron Shaw Fordyce, Philipp Koehn, Christof Monz, and Josh Schroeder. 2007 · 2007
Earlier work this paper cites.
The NIST 2008 Metrics for machine translation challenge—overview, methodology, metrics, and results
Mark Przybocki, Kay Peterson, Sébastien Bronsart, and Gregory Sanders. 2009 · 2008
Earlier work this paper cites.
Continuous measurement scales in human evaluation of machine translation
Yvette Graham, Timothy Baldwin, Alistair Moffat, and Justin Zobel. 2013 · 2013
Earlier work this paper cites.
Results of the WMT13 Metrics Shared Task
Matouš Macháček and Ondřej Bojar. 2013 · 2013
Earlier work this paper cites.
Results of the WMT14 Metrics Shared Task
Matouš Macháček and Ondřej Bojar. 2014 · 2014
Cited alongside, same era.
chrF: character n-gram F-score for automatic MT evaluation
Maja Popović. 2015 · 2015
Cited alongside, same era.
Results of the WMT15 Metrics Shared Task
Miloš Stanojević, Amir Kamran, Philipp Koehn, and Ondřej Bojar. 2015 · 2015
Cited alongside, same era.
Findings of the 2016 conference on machine translation
Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Aurelie Neveol, Mariana Neves, Martin Popel, Matt Post, Raphael Rubino, Carolina Scarton, Lucia Specia, Marco Turchi, Karin Verspoor, and Marcos Zampieri. 2016a · 2016
Cited alongside, same era.
CharacTer: Translation edit rate on character level
Weiyue Wang, Jan-Thorsten Peter, Hendrik Rosendahl, and Hermann Ney. 2016 · 2016
Cited alongside, same era.
Results of the WMT17 Metrics Shared Task
Putting evaluation in context: Contextual embeddings improve machine translation evaluation
Nitika Mathur, Timothy Baldwin, and Trevor Cohn. 2019 · 2019
Later among the works it cites.
Fluent translations from disfluent speech in end-to-end speech translation
Elizabeth Salesky, Matthias Sperber, and Alexander Waibel. 2019 · 2019
Later among the works it cites.
EED: Extended Edit Distance Measure for Machine Translation
Peter Stanchev, Weiyue Wang, and Hermann Ney. 2019 · 2019
Later among the works it cites.
Moving to a world beyond "p<0.05"
Ronald L. Wasserstein, Allen L. Schirm, and Nicole A. Lazar. 2019 · 2019
Later among the works it cites.
The effect of translationese in machine translation test sets
Mike Zhang and Antonio Toral. 2019 · 2019
Later among the works it cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ondřej Bojar, Yvette Graham, and Amir Kamran. 2017 · 2017
Cited alongside, same era.
Can machine translation systems be evaluated by the crowd alone
Yvette Graham, Timothy Baldwin, Alistair Moffat, and Justin Zobel. 2017 · 2017
Cited alongside, same era.
The hitchhiker’s guide to testing statistical significance in natural language processing
Rotem Dror, Gili Baumer, Segev Shlomov, and Roi Reichart. 2018 · 2018
Cited alongside, same era.
Appraise evaluation framework for machine translation
Christian Federmann. 2018 · 2018
Cited alongside, same era.
Results of the WMT18 metrics shared task: Both characters and embeddings achieve good performance
Qingsong Ma, Ondřej Bojar, and Yvette Graham. 2018 · 2018
Cited alongside, same era.
A Call for Clarity in Reporting BLEU Scores
Matt Post. 2018 · 2018
Cited alongside, same era.
A structured review of the validity of bleu
Ehud Reiter. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
BLEU might be guilty but references are not innocent
Markus Freitag, David Grangier, and Isaac Caswell. 2020 · 2020
Later among the works it cites.
Statistical power and translationese in machine translation evaluation
Yvette Graham, Barry Haddow, and Philipp Koehn. 2020 · 2020
Later among the works it cites.
COMET: A Neural Framework for MT Evaluation
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020 · 2020
Later among the works it cites.
BLEURT: Learning Robust Metrics for Text Generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Later among the works it cites.
Automatic Machine Translation Evaluation in Many Languages via Zero-Shot Paraphrasing
Brian Thompson and Matt Post. 2020 · 2020
Later among the works it cites.
Assessing reference-free peer evaluation for machine translation
Sweta Agrawal, George Foster, Markus Freitag, and Colin Cherry. 2021 · 2021
Closest in time.
Experts, errors, and context: A large-scale study of human evaluation for machine translation
Markus Freitag, George Foster, David Grangier, Viresh Ratnakar, Qijun Tan, and Wolfgang Macherey. 2021 · 2021
Closest in time.
Scientific credibility of machine translation research: A meta-evaluation of 769 papers
Benjamin Marie, Atsushi Fujita, and Raphael Rubino. 2021 · 2021
Closest in time.