Fetching the paper…
Reading the bibliography…
Ten years ago a single metric, BLEU, governed progress in machine translation research.
A method for the solution of certain non-linear problems in least squares
Kenneth Levenberg. 1944 · 1944
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Holy and unholy grails
Eduard Hovy and Deepak Ravichandran. 2003 · 2003
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
BLEU in characters: Towards automatic MT evaluation in languages without word delimiters
Etienne Denoual and Yves Lepage. 2005 · 2005
Earlier work this paper cites.
Comparing automatic and human evaluation of NLG systems
Anja Belz and Ehud Reiter. 2006 · 2006
Earlier work this paper cites.
Re-evaluating the role of Bleu in machine translation research
Chris Callison-Burch, Miles Osborne, and Philipp Koehn. 2006 · 2006
Earlier work this paper cites.
Finding software metrics threshold values using ROC curves
Raed Shatnawi, Wei Li, James Swain, and Tim Newman. 2010 · 2010
Earlier work this paper cites.
An empirical investigation of statistical significance in NLP
Taylor Berg-Kirkpatrick, David Burkett, and Dan Klein. 2012 · 2012
Earlier work this paper cites.
On effect size
Ken Kelley and Kristopher J Preacher. 2012 · 2012
Earlier work this paper cites.
How big is “big”? interpreting effect sizes in l2 research
Luke Plonsky and Frederick L Oswald. 2014 · 2014
Earlier work this paper cites.
chrF: character n-gram F-score for automatic MT evaluation
Maja Popović. 2015 · 2015
Earlier work this paper cites.
Findings of the 2016 conference on machine translation
Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Aurélie Névéol, Mariana Neves, Martin Popel, Matt Post, Raphael Rubino, Carolina Scarton, Lucia Specia, Marco Turchi, Karin Verspoor, and Marcos Zampieri. 2016 · 2016
Cited alongside, same era.
Statistical tests, p values, confidence intervals, and power: A guide to misinterpretations
Sander Greenland, Stephen J Senn, Kenneth J Rothman, John B Carlin, Charles Poole, Steven N Goodman, and Douglas G Altman. 2016 · 2016
Cited alongside, same era.
Evaluating domain-specific metric thresholds: An empirical study
Allan Mori, Gustavo Vale, Markos Viggiato, Johnatan Oliveira, Eduardo Figueiredo, Elder Cirilo, Pooyan Jamshidi, and Christian Kastner. 2018 · 2018
Cited alongside, same era.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Cited alongside, same era.
Attaining the unattainable? reassessing claims of human parity in neural machine translation
Antonio Toral, Sheila Castilho, Ke Hu, and Andy Way. 2018 · 2018
ACES: Translation accuracy challenge sets for evaluating machine translation metrics
Chantal Amrhein, Nikita Moghe, and Liane Guillou. 2022 · 2022
Later among the works it cites.
The Flores-101 evaluation benchmark for low-resource and multilingual machine translation
Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc’Aurelio Ranzato, Francisco Guzmán, and Angela Fan. 2022 · 2022
Later among the works it cites.
DEMETR: Diagnosing evaluation metrics for translation
Marzena Karpinska, Nishant Raj, Katherine Thai, Yixiao Song, Ankita Gupta, and Mohit Iyyer. 2022 · 2022
Later among the works it cites.
Findings of the 2022 conference on machine translation (WMT22)
Tom Kocmi, Rachel Bawden, Ondřej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Thamme Gowda, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Rebecca Knowles, Philipp Koehn, Christof Monz, Makoto Morishita, Masaaki Nagata, Toshiaki Nakazawa, Michal Novák, Martin Popel, and Maja Popović. 2022 · 2022
Later among the works it cites.
Yes, we need statistical significance testing
Benjamin Marie. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Cited alongside, same era.
COMET: A neural framework for MT evaluation
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020 · 2020
Cited alongside, same era.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Cited alongside, same era.
To ship or not to ship: An extensive evaluation of automatic metrics for machine translation
Tom Kocmi, Christian Federmann, Roman Grundkiewicz, Marcin Junczys-Dowmunt, Hitokazu Matsushita, and Arul Menezes. 2021 · 2021
Cited alongside, same era.
Scientific credibility of machine translation research: A meta-evaluation of 769 papers
Benjamin Marie, Atsushi Fujita, and Raphael Rubino. 2021 · 2021
Cited alongside, same era.
Learning compact metrics for MT
Amy Pu, Hyung Won Chung, Ankur Parikh, Sebastian Gehrmann, and Thibault Sellam. 2021 · 2021
Cited alongside, same era.
The devil is in the errors: Leveraging large language models for fine-grained machine translation evaluation
Patrick Fernandes, Daniel Deutsch, Mara Finkelstein, Parker Riley, André Martins, Graham Neubig, Ankush Garg, Jonathan Clark, Markus Freitag, and Orhan Firat. 2023 · 2023
Later among the works it cites.
Results of WMT23 metrics shared task: Metrics might be guilty but references are not innocent
Markus Freitag, Nitika Mathur, Chi-kiu Lo, Eleftherios Avramidis, Ricardo Rei, Brian Thompson, Tom Kocmi, Frederic Blain, Daniel Deutsch, Craig Stewart, Chrysoula Zerva, Sheila Castilho, Alon Lavie, and George Foster. 2023 · 2023
Later among the works it cites.
GEMBA-MQM: Detecting translation quality error spans with GPT-4
Tom Kocmi and Christian Federmann. 2023 · 2023
Later among the works it cites.
Llms as narcissistic evaluators: When ego inflates evaluation scores
Yiqi Liu, Nafise Sadat Moosavi, and Chenghua Lin. 2023 · 2023
Later among the works it cites.
Beyond correlation: Making sense of the score differences of new MT evaluation metrics
Chi-kiu Lo, Rebecca Knowles, and Cyril Goutte. 2023 · 2023
Later among the works it cites.
There’s no data like better data: Using QE metrics for MT data filtering
Jan-Thorsten Peter, David Vilar, Daniel Deutsch, Mara Finkelstein, Juraj Juraska, and Markus Freitag. 2023 · 2023
Later among the works it cites.
Quality and quantity of machine translation references for automated metrics
Vilém Zouhar and Ondřej Bojar. 2024 · 2024
Closest in time.