Fetching the paper…
Reading the bibliography…
This paper presents the first large-scale meta-evaluation of machine translation (MT).
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Statistical significance tests for machine translation evaluation
Philipp Koehn. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
On some pitfalls in automatic evaluation and significance testing for MT
Stefan Riezler and John T. Maxwell. 2005 · 2005
Earlier work this paper cites.
Re-evaluating the role of Bleu in machine translation research
Chris Callison-Burch, Miles Osborne, and Philipp Koehn. 2006 · 2006
Earlier work this paper cites.
A study of translation edit rate with targeted human annotation
Matthew Snover, Bonnie Dorr, Richard Schwartz, Linnea Micciulla, and John Makhoul. 2006 · 2006
Earlier work this paper cites.
Automatic evaluation of translation quality for distant language pairs
Hideki Isozaki, Tsutomu Hirao, Kevin Duh, Katsuhito Sudoh, and Hajime Tsukada. 2010 · 2010
Earlier work this paper cites.
chrF: character n-gram F-score for automatic MT evaluation
Maja Popović. 2015 · 2015
Cited alongside, same era.
Stronger baselines for trustable results in neural machine translation
Michael Denkowski and Graham Neubig. 2017 · 2017
Cited alongside, same era.
The hitchhiker’s guide to testing statistical significance in natural language processing
Rotem Dror, Gili Baumer, Segev Shlomov, and Roi Reichart. 2018 · 2018
Cited alongside, same era.
Marian: Fast neural machine translation in C++
Marcin Junczys-Dowmunt, Roman Grundkiewicz, Tomasz Dwojak, Hieu Hoang, Kenneth Heafield, Tom Neckermann, Frank Seide, Ulrich Germann, Alham Fikri Aji, Nikolay Bogoychev, André F. T. Martins, and Alexandra Birch. 2018 · 2018
Cited alongside, same era.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Cited alongside, same era.
A structured review of the validity of BLEU
Ehud Reiter. 2018 · 2018
Moving to a world beyond “ p p < < 0.05”
Ronald L. Wasserstein, Allen L. Schirm, and Nicole A. Lazar. 2019 · 2019
Later among the works it cites.
Findings of the 2020 Conference on Machine Translation (WMT20)
Loïc Barrault, Magdalena Biesialska, Ondřej Bojar, Marta R. Costa-jussà, Christian Federmann, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Matthias Huck, Eric Joanis, Tom Kocmi, Philipp Koehn, Chi-kiu Lo, Nikola Ljubešić, Christof Monz, Makoto Morishita, Masaaki Nagata, Toshiaki Nakazawa, Santanu Pal, Matt Post, and Marcos Zampieri. 2020 · 2020
Later among the works it cites.
Dynamic context selection for document-level neural machine translation via reinforcement learning
Xiaomian Kang, Yang Zhao, Jiajun Zhang, and Chengqing Zong. 2020 · 2020
Later among the works it cites.
A set of recommendations for assessing human–machine parity in language translation
Samuel Läubli, Sheila Castilho, Graham Neubig, Rico Sennrich, Qinlan Shen, and Antonio Toral. 2020 · 2020
Later among the works it cites.
A simple and effective unified encoder for document-level machine translation
Shuming Ma, Dongdong Zhang, and Ming Zhou. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Selective attention for context-aware neural machine translation
Sameen Maruf, André F. T. Martins, and Gholamreza Haffari. 2019 · 2019
Cited alongside, same era.
Later among the works it cites.
Tangled up in BLEU: Reevaluating the evaluation of automatic machine translation evaluation metrics
Nitika Mathur, Timothy Baldwin, and Trevor Cohn. 2020 · 2020
Later among the works it cites.