Fetching the paper…
Reading the bibliography…
Automatic metrics are fundamental for the development and evaluation of machine translation systems.
Graphs in statistical analysis
Francis J Anscombe. 1973 · 1973
Earlier work this paper cites.
Outliers in Statistical Data
Vic Barnett and Toby Lewis. 1974 · 1974
Earlier work this paper cites.
How to detect and handle outliers , volume 16
Boris Iglewicz and David Caster Hoaglin. 1993 · 1993
Earlier work this paper cites.
Statistical significance tests for machine translation evaluation
Philipp Koehn. 2004 · 2004
Earlier work this paper cites.
The power of outliers (and why researchers should always check for them)
Jason W Osborne and Amy Overbay. 2004 · 2004
Earlier work this paper cites.
Inferences based on a skipped correlation coefficient
Rand Wilcox. 2004 · 2004
Earlier work this paper cites.
Evaluating evaluation methods for generation in the presence of variation
Amanda Stent, Matthew Marge, and Mohit Singhai. 2005 · 2005
Earlier work this paper cites.
Re-evaluating the role of Bleu in machine translation research
Chris Callison-Burch, Miles Osborne, and Philipp Koehn. 2006 · 2006
Earlier work this paper cites.
A study of translation edit rate with targeted human annotation
Matthew Snover, Bonnie Dorr, Richard Schwartz, Linnea Micciulla, and John Makhoul. 2006 · 2006
Earlier work this paper cites.
Robust statistics for outlier detection
Peter J Rousseeuw and Mia Hubert. 2011 · 2011
Cited alongside, same era.
Detecting outliers: Do not use standard deviation around the mean, use absolute deviation around the median
Christophe Leys, Christophe Ley, Olivier Klein, Philippe Bernard, and Laurent Licata. 2013 · 2013
Cited alongside, same era.
Findings of the 2014 workshop on statistical machine translation
Ondřej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, Christof Monz, Pavel Pecina, Matt Post, Herve Saint-Amand, Radu Soricut, Lucia Specia, and Aleš Tamchyna. 2014 · 2014
Cited alongside, same era.
Testing for significance of increased correlation with human judgment
Yvette Graham and Timothy Baldwin. 2014 · 2014
Cited alongside, same era.
Randomized significance tests in machine translation
Yvette Graham, Nitika Mathur, and Timothy Baldwin. 2014 · 2014
Cited alongside, same era.
Can machine translation systems be evaluated by the crowd alone?
Yvette Graham, Timothy Baldwin, Alistair Moffat, and Justin Zobel. 2017 · 2017
Later among the works it cites.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Later among the works it cites.
A structured review of the validity of BLEU
Ehud Reiter. 2018 · 2018
Later among the works it cites.
Findings of the 2019 conference on machine translation (WMT19)
Loïc Barrault, Ondřej Bojar, Marta R. Costa-jussà, Christian Federmann, Mark Fishel, Yvette Graham, Barry Haddow, Matthias Huck, Philipp Koehn, Shervin Malmasi, Christof Monz, Mathias Müller, Santanu Pal, Matt Post, and Marcos Zampieri. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Taking MT evaluation metrics to extremes: Beyond correlation with human judgments
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Maja Popović. 2015 · 2015
Cited alongside, same era.
Findings of the 2016 conference on machine translation
Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Aurélie Névéol, Mariana Neves, Martin Popel, Matt Post, Raphael Rubino, Carolina Scarton, Lucia Specia, Marco Turchi, Karin Verspoor, and Marcos Zampieri. 2016 · 2016
Cited alongside, same era.
Climbing mont BLEU: The strange world of reachable high-BLEU translations
Aaron Smith, Christian Hardmeier, and Joerg Tiedemann. 2016 · 2016
Cited alongside, same era.
Enhanced LSTM for natural language inference
Qian Chen, Xiaodan Zhu, Zhen-Hua Ling, Si Wei, Hui Jiang, and Diana Inkpen. 2017 · 2017
Cited alongside, same era.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002a
Cited in the paper.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002b
Cited in the paper.
Marina Fomicheva and Lucia Specia. 2019 · 2019
Later among the works it cites.
YiSi — a unified semantic MT quality evaluation and estimation metric for languages with different levels of available resources
Chi-kiu Lo. 2019 · 2019
Later among the works it cites.
Results of the WMT19 metrics shared task: Segment-level and strong MT systems pose big challenges
Qingsong Ma, Johnny Wei, Ondřej Bojar, and Yvette Graham. 2019 · 2019
Later among the works it cites.
Putting evaluation in context: Contextual embeddings improve machine translation evaluation
Nitika Mathur, Timothy Baldwin, and Trevor Cohn. 2019 · 2019
Later among the works it cites.