Fetching the paper…
Reading the bibliography…
Kendall's tau is frequently used to meta-evaluate how well machine translation (MT) evaluation metrics score individual translations.
Language Models are Few-Shot Learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
A New Measure of Rank Correlation
Maurice G Kendall. 1938 · 1938
Earlier work this paper cites.
The Treatment of Ties in Ranking Problems
Maurice G Kendall. 1945 · 1945
Earlier work this paper cites.
The Estimation and Comparison of Strengths of Association in Contingency Tables
Alan Stuart. 1953 · 1953
Earlier work this paper cites.
Evaluating Machine Translation Output with Automatic Sentence Segmentation
Evgeny Matusov, Gregor Leusch, Oliver Bender, and Hermann Ney. 2005 · 2005
Earlier work this paper cites.
Findings of the 2010 Joint Workshop on Statistical Machine Translation and Metrics for Machine Translation
Chris Callison-Burch, Philipp Koehn, Christof Monz, Kay Peterson, Mark Przybocki, and Omar Zaidan. 2010 · 2010
Earlier work this paper cites.
Results of the WMT13 Metrics Shared Task
Matouš Macháček and Ondřej Bojar. 2013 · 2013
Earlier work this paper cites.
Multidimensional Quality Metrics (MQM): A Framework for Declaring and Describing Translation Quality Metrics
Arle Lommel, Hans Uszkoreit, and Aljoscha Burchardt. 2014 · 2014
Cited alongside, same era.
Results of the WMT14 Metrics Shared Task
Matouš Macháček and Ondřej Bojar. 2014 · 2014
Cited alongside, same era.
Results of the WMT17 Metrics Shared Task
Ondřej Bojar, Yvette Graham, and Amir Kamran. 2017 · 2017
Cited alongside, same era.
Results of the WMT18 Metrics Shared Task: Both characters and embeddings achieve good performance
Qingsong Ma, Ondřej Bojar, and Yvette Graham. 2018 · 2018
Cited alongside, same era.
Tangled up in bleu: Reevaluating the evaluation of automatic machine translation evaluation metrics
Nitika Mathur, Timothy Baldwin, and Trevor Cohn. 2020 · 2020
Cited alongside, same era.
COMET: A Neural Framework for MT Evaluation
BLEURT: Learning Robust Metrics for Text Generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Later among the works it cites.
To Ship or Not to Ship: An Extensive Evaluation of Automatic Metrics for Machine Translation
Tom Kocmi, Christian Federmann, Roman Grundkiewicz, Marcin Junczys-Dowmunt, Hitokazu Matsushita, and Arul Menezes. 2021 · 2021
Later among the works it cites.
Quality-Aware Decoding for Neural Machine Translation
Patrick Fernandes, António Farinhas, Ricardo Rei, José G. C. de Souza, Perez Ogayo, Graham Neubig, and Andre Martins. 2022 · 2022
Later among the works it cites.
MaTESe: Machine Translation Evaluation as a Sequence Tagging Problem
Stefano Perrella, Lorenzo Proietti, Alessandro Scirè, Niccolò Campolungo, and Roberto Navigli. 2022 · 2022
Later among the works it cites.
COMET-22: Unbabel-IST 2022 Submission for the Metrics Shared Task
Ricardo Rei, José G. C. de Souza, Duarte Alves, Chrysoula Zerva, Ana C Farinha, Taisiya Glushkova, Alon Lavie, Luisa Coheur, and André F. T. Martins. 2022 · 2022
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020 · 2020
Cited alongside, same era.
Experts, Errors, and Context: A Large-Scale Study of Human Evaluation for Machine Translation
Markus Freitag, George Foster, David Grangier, Viresh Ratnakar, Qijun Tan, and Wolfgang Macherey. 2021a
Cited in the paper.
High Quality Rather than High Model Probability: Minimum Bayes Risk Decoding with Neural Metrics
Markus Freitag, David Grangier, Qijun Tan, and Bowen Liang. 2022a
Cited in the paper.
Results of WMT22 Metrics Shared Task: Stop Using BLEU – Neural Metrics Are Better and More Robust
Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, Eleftherios Avramidis, Tom Kocmi, George Foster, Alon Lavie, and André F. T. Martins. 2022b
Cited in the paper.
Results of the WMT21 Metrics Shared Task: Evaluating Metrics with Expert-based Human Evaluations on TED and News Domain
Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, George Foster, Alon Lavie, and Ondřej Bojar. 2021b
Cited in the paper.
Large Language Models Are State-of-the-Art Evaluators of Translation Quality
Tom Kocmi and Christian Federmann. 2023 · 2023
Closest in time.