Fetching the paper…
Reading the bibliography…
Human evaluation of modern high-quality machine translation systems is a difficult problem, and there is increasing evidence that inadequate evaluation procedures can lead to erroneous conclusions.
Language and Machines: Computers in Translation and Linguistics; a Report , volume 1416
ALPAC. 1966 · 1966
Earlier work this paper cites.
The arpa mt evaluation methodologies: evolution, lessons, and future approaches
John S White, Theresa A O’Connell, and Francis E O’Mara. 1994 · 1994
Earlier work this paper cites.
Manual and Automatic Evaluation of Machine Translation between European Languages
Philipp Koehn and Christof Monz. 2006 · 2006
Earlier work this paper cites.
Human Evaluation of Machine Translation Through Binary System Comparisons
David Vilar, Gregor Leusch, Hermann Ney, and Rafael E Banchs. 2007 · 2007
Earlier work this paper cites.
Proceedings of the Third Workshop on Statistical Machine Translation
Chris Callison-Burch, Philipp Koehn, Christof Monz, Josh Schroeder, and Cameron Shaw Fordyce. 2008 · 2008
Earlier work this paper cites.
Translationese and Its Dialects
Moshe Koppel and Noam Ordan. 2011 · 2011
Earlier work this paper cites.
Involving Language Professionals in the Evaluation of Machine Translation
Eleftherios Avramidis, Aljoscha Burchardt, Christian Federmann, Maja Popović, Cindy Tscherwinka, and David Vilar. 2012 · 2012
Earlier work this paper cites.
Continuous Measurement Scales in Human Evaluation of Machine Translation
Yvette Graham, Timothy Baldwin, Alistair Moffat, and Justin Zobel. 2013 · 2013
Earlier work this paper cites.
Multidimensional Quality Metrics (MQM) : A Framework for Declaring and Describing Translation Quality Metrics
Arle Lommel, Hans Uszkoreit, and Aljoscha Burchardt. 2014 · 2014
Earlier work this paper cites.
Findings of the 2016 Conference on Machine Translation
Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Aurélie Névéol, Mariana Neves, Martin Popel, Matt Post, Raphael Rubino, Carolina Scarton, Lucia Specia, Marco Turchi, Karin Verspoor, and Marcos Zampieri. 2016 · 2016
Earlier work this paper cites.
A Reading Comprehension Corpus for Machine Translation Evaluation
Carolina Scarton and Lucia Specia. 2016 · 2016
Earlier work this paper cites.
Findings of the 2017 Conference on Machine Translation (WMT17)
Ondrej Bojar, Rajen Chatterjee, Federmann Christian, Graham Yvette, Haddow Barry, Huck Matthias, Koehn Philipp, Liu Qun, Logacheva Varvara, Monz Christof, et al. 2017 · 2017
Cited alongside, same era.
A Comparative Quality Evaluation of PBSMT and NMT using Professional Translators
Sheila Castilho, Joss Moorkens, Federico Gaspari, Rico Sennrich, Vilelmini Sosoni, Panayota Georgakopoulou, Pintu Lohar, Andy Way, Antonio Valerio Miceli Barone, and Maria Gialama. 2017 · 2017
Cited alongside, same era.
The Role of Human Reference Translation in Machine Translation Evaluation
Marina Fomicheva et al. 2017 · 2017
Cited alongside, same era.
Can Machine Translation Systems be Evaluated by the Crowd Alone?
Yvette Graham, Timothy Baldwin, Alistair Moffat, and Justin Zobel. 2017 · 2017
Cited alongside, same era.
Machine Translation Human Evaluation: an investigation of evaluation based on Post-Editing and its relation with Direct Assessment
Luisa Bentivogli, Mauro Cettolo, Marcello Federico, and Christian Federmann. 2018 · 2018
Cited alongside, same era.
APE at Scale and Its Implications on MT Evaluation Biases
Markus Freitag, Isaac Caswell, and Scott Roy. 2019 · 2019
Later among the works it cites.
Reassessing claims of human parity and super-human performance in machine translation at wmt 2019
Antonio Toral. 2020 · 2019
Later among the works it cites.
The Effect of Translationese in Machine Translation Test Sets
Mike Zhang and Antonio Toral. 2019 · 2019
Later among the works it cites.
Findings of the 2020 Conference on Machine Translation (WMT20)
Loïc Barrault, Magdalena Biesialska, Ondřej Bojar, Marta R. Costa-jussà, Christian Federmann, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Matthias Huck, Eric Joanis, Tom Kocmi, Philipp Koehn, Chi-kiu Lo, Nikola Ljubešić, Christof Monz, Makoto Morishita, Masaaki Nagata, Toshiaki Nakazawa, Santanu Pal, Matt Post, and Marcos Zampieri. 2020 · 2020
Later among the works it cites.
What’s the Difference Between Professional Human and Machine Translation? A Blind Multi-language Study on Domain-specific MT
Lukas Fischer and Samuel Läubli. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Exploring Gap Filling as a Cheaper Alternative to Reading Comprehension Questionnaires when Evaluating Machine Translation for Gisting
Mikel L Forcada, Carolina Scarton, Lucia Specia, Barry Haddow, and Alexandra Birch. 2018 · 2018
Cited alongside, same era.
Achieving Human Parity on Automatic Chinese to English News Translation
Hany Hassan, Anthony Aue, Chang Chen, Vishal Chowdhary, Jonathan Clark, Christian Federmann, Xuedong Huang, Marcin Junczys-Dowmunt, William Lewis, Mu Li, et al. 2018 · 2018
Cited alongside, same era.
Quantitative Fine-Grained Human Evaluation of Machine Translation Systems: a Case Study on English to Croatian
Filip Klubička, Antonio Toral, and Víctor M Sánchez-Cartagena. 2018 · 2018
Cited alongside, same era.
Has Machine Translation Achieved Human Parity? A Case for Document-level Evaluation
Samuel Läubli, Rico Sennrich, and Martin Volk. 2018 · 2018
Cited alongside, same era.
Comparing Bayesian Models of Annotation
Silviu Paun, Bob Carpenter, Jon Chamberlain, Dirk Hovy, Udo Kruschwitz, and Massimo Poesio. 2018 · 2018
Cited alongside, same era.
Attaining the Unattainable? Reassessing Claims of Human Parity in Neural Machine Translation
Antonio Toral, Sheila Castilho, Ke Hu, and Andy Way. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
BLEU might be Guilty but References Are Not Innocent
Markus Freitag, David Grangier, and Isaac Caswell. 2020 · 2020
Later among the works it cites.
Translationese in Machine Translation Evaluation
Yvette Graham, Barry Haddow, and Philipp Koehn. 2020 · 2020
Later among the works it cites.
A Set of Recommendations for Assessing Human–Machine Parity in Language Translation
Samuel Läubli, Sheila Castilho, Graham Neubig, Rico Sennrich, Qinlan Shen, and Antonio Toral. 2020 · 2020
Later among the works it cites.
Results of the WMT20 Metrics Shared Task
Nitika Mathur, Johnny Wei, Markus Freitag, Qingsong Ma, and Ondřej Bojar. 2020 · 2020
Later among the works it cites.
Informative Manual Evaluation of Machine Translation Output
Maja Popović. 2020 · 2020
Later among the works it cites.
COMET: A Neural Framework for MT Evaluation
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020 · 2020
Later among the works it cites.