Fetching the paper…
Reading the bibliography…
High-quality Machine Translation (MT) evaluation relies heavily on human judgments.
Language and machines. Computers in translation and linguistics
ALPAC. 1966 · 1966
Earlier work this paper cites.
The ARPA MT evaluation methodologies: Evolution, lessons, and future approaches
John S. White, Theresa A. O’Connell, and Francis E. O’Mara. 1994 · 1994
Earlier work this paper cites.
Manual and automatic evaluation of machine translation between European languages
Philipp Koehn and Christof Monz. 2006 · 2006
Earlier work this paper cites.
Error analysis of statistical machine translation output
David Vilar, Jia Xu, Luis Fernando D’Haro, and Hermann Ney. 2006 · 2006
Earlier work this paper cites.
(meta-) evaluation of machine translation
Chris Callison-Burch, Cameron Fordyce, Philipp Koehn, Christof Monz, and Josh Schroeder. 2007 · 2007
Earlier work this paper cites.
Human evaluation of machine translation through binary system comparisons
David Vilar, Gregor Leusch, Hermann Ney, and Rafael E. Banchs. 2007 · 2007
Earlier work this paper cites.
Further meta-evaluation of machine translation
Chris Callison-Burch, Cameron Fordyce, Philipp Koehn, Christof Monz, and Josh Schroeder. 2008 · 2008
Earlier work this paper cites.
Multidimensional quality metrics: a flexible system for assessing translation quality
Aljoscha Burchardt. 2013 · 2013
Earlier work this paper cites.
Continuous measurement scales in human evaluation of machine translation
Yvette Graham, Timothy Baldwin, Alistair Moffat, and Justin Zobel. 2013 · 2013
Earlier work this paper cites.
Using a new analytic measure for the annotation and analysis of MT errors on real data
Arle Lommel, Aljoscha Burchardt, Maja Popović, Kim Harris, Eleftherios Avramidis, and Hans Uszkoreit. 2014 · 2014
Cited alongside, same era.
Findings of the 2016 conference on machine translation
Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Aurélie Névéol, Mariana Neves, Martin Popel, Matt Post, Raphael Rubino, Carolina Scarton, Lucia Specia, Marco Turchi, Karin Verspoor, and Marcos Zampieri. 2016 · 2016
Cited alongside, same era.
Findings of the 2017 conference on machine translation (WMT17)
Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Shujian Huang, Matthias Huck, Philipp Koehn, Qun Liu, Varvara Logacheva, Christof Monz, Matteo Negri, Matt Post, Raphael Rubino, Lucia Specia, and Marco Turchi. 2017 · 2017
Cited alongside, same era.
Appraise evaluation framework for machine translation
Christian Federmann. 2018 · 2018
Cited alongside, same era.
Quantitative Fine-grained Human Evaluation of Machine Translation Systems: A Case Study on English to Croatian
To ship or not to ship: An extensive evaluation of automatic metrics for machine translation
Tom Kocmi, Christian Federmann, Roman Grundkiewicz, Marcin Junczys-Dowmunt, Hitokazu Matsushita, and Arul Menezes. 2021 · 2021
Later among the works it cites.
Findings of the 2022 conference on machine translation (WMT22)
Tom Kocmi, Rachel Bawden, Ondřej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Thamme Gowda, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Rebecca Knowles, Philipp Koehn, Christof Monz, Makoto Morishita, Masaaki Nagata, Toshiaki Nakazawa, Michal Novák, Martin Popel, and Maja Popović. 2022 · 2022
Later among the works it cites.
Two contrasting data annotation paradigms for subjective NLP tasks
Paul Rottger, Bertie Vidgen, Dirk Hovy, and Janet Pierrehumbert. 2022 · 2022
Later among the works it cites.
Searching for a higher power in the human evaluation of MT
Johnny Wei, Tom Kocmi, and Christian Federmann. 2022 · 2022
Later among the works it cites.
Results of WMT23 metrics shared task: Metrics might be guilty but references are not innocent
Markus Freitag, Nitika Mathur, Chi-kiu Lo, Eleftherios Avramidis, Ricardo Rei, Brian Thompson, Tom Kocmi, Frederic Blain, Daniel Deutsch, Craig Stewart, Chrysoula Zerva, Sheila Castilho, Alon Lavie, and George Foster. 2023 · 2023
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Filip Klubička, Antonio Toral, and Víctor M. Sánchez-Cartagena. 2018 · 2018
Cited alongside, same era.
Fluency over adequacy: A pilot study in measuring user trust in imperfect MT
Marianna Martindale and Marine Carpuat. 2018 · 2018
Cited alongside, same era.
On the same page? comparing inter-annotator agreement in sentence and document level human machine translation evaluation
Sheila Castilho. 2020 · 2020
Cited alongside, same era.
Correct me if you can: Learning from error corrections and markings
Julia Kreutzer, Nathaniel Berger, and Stefan Riezler. 2020 · 2020
Cited alongside, same era.
Informative manual evaluation of machine translation output
Maja Popović. 2020 · 2020
Cited alongside, same era.
Experts, errors, and context: A large-scale study of human evaluation for machine translation
Markus Freitag, George Foster, David Grangier, Viresh Ratnakar, Qijun Tan, and Wolfgang Macherey. 2021a
Cited in the paper.
Results of the WMT21 metrics shared task: Evaluating metrics with expert-based human evaluations on TED and news domain
Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, George Foster, Alon Lavie, and Ondřej Bojar. 2021b
Cited in the paper.
Later among the works it cites.
Findings of the 2023 conference on machine translation (WMT23): LLMs are here but not quite there yet
Tom Kocmi, Eleftherios Avramidis, Rachel Bawden, Ondřej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Markus Freitag, Thamme Gowda, Roman Grundkiewicz, Barry Haddow, Philipp Koehn, Benjamin Marie, Christof Monz, Makoto Morishita, Kenton Murray, Makoto Nagata, Toshiaki Nakazawa, Martin Popel, Maja Popović, and Mariya Shmatova. 2023 · 2023
Later among the works it cites.
Calibration and context in human evaluation of machine translation
Rebecca Knowles and Chi-kiu Lo. 2024 · 2024
Closest in time.
Finding replicable human evaluations via stable ranking probability
Parker Riley, Daniel Deutsch, George Foster, Viresh Ratnakar, Ali Dabirmoghaddam, and Markus Freitag. 2024 · 2024
Closest in time.
AI-assisted human evaluation of machine translation
Vilém Zouhar, Tom Kocmi, and Mrinmaya Sachan. 2024 · 2024
Closest in time.