Fetching the paper…
Reading the bibliography…
Automated metrics for Machine Translation have made significant progress, with the goal of replacing expensive and time-consuming human evaluations.
The isotonic regression problem and its dual
R. E. Barlow and H. D. Brunk. 1972 · 1972
Earlier work this paper cites.
Comparing rating scales and preference judgements in language evaluation
Anja Belz and Eric Kow. 2010 · 2010
Earlier work this paper cites.
Scikit-learn: Machine learning in Python
F. Pedregosa, G. Varoquaux, A. Gramfort, V. Michel, B. Thirion, O. Grisel, M. Blondel, P. Prettenhofer, R. Weiss, V. Dubourg, J. Vanderplas, A. Passos, D. Cournapeau, M. Brucher, M. Perrot, and E. Duchesnay. 2011 · 2011
Earlier work this paper cites.
Multidimensional quality metrics (mqm) : a framework for declaring and describing translation quality metrics
Arle. Language Technology Lab) Lommel, Hans. Language Technology Lab) Uszkoreit, and Aljoscha. Language Technology Lab) Burchardt. 2014 · 2014
Earlier work this paper cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
The price of debiasing automatic metrics in natural language evalaution
Arun Chaganty, Stephen Mussmann, and Percy Liang. 2018 · 2018
Earlier work this paper cites.
Language models are few-shot learners
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020 · 2020
Earlier work this paper cites.
Array programming with NumPy
Charles R. Harris, K. Jarrod Millman, Stéfan J. van der Walt, Ralf Gommers, Pauli Virtanen, David Cournapeau, Eric Wieser, Julian Taylor, Sebastian Berg, Nathaniel J. Smith, Robert Kern, Matti Picus, Stephan Hoyer, Marten H. van Kerkwijk, Matthew Brett, Allan Haldane, Jaime Fernández del Río, Mark Wiebe, Pearu Peterson, Pierre Gérard-Marchant, Kevin Sheppard, Tyler Reddy, Warren Weckesser, Hameer Abbasi, Christoph Gohlke, and Travis E. Oliphant. 2020 · 2020
Earlier work this paper cites.
Automatic machine translation evaluation in many languages via zero-shot paraphrasing
Brian Thompson and Matt Post. 2020a · 2020
Earlier work this paper cites.
Results of the WMT21 metrics shared task: Evaluating metrics with expert-based human evaluations on TED and news domain
Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, George Foster, Alon Lavie, and Ondřej Bojar. 2021 · 2021
Cited alongside, same era.
To ship or not to ship: An extensive evaluation of automatic metrics for machine translation
Tom Kocmi, Christian Federmann, Roman Grundkiewicz, Marcin Junczys-Dowmunt, Hitokazu Matsushita, and Arul Menezes. 2021 · 2021
Cited alongside, same era.
Estimating expected calibration errors
Nicolas Posocco and Antoine Bonnefoy. 2021 · 2021
Cited alongside, same era.
The statistical advantage of automatic NLG metrics at the system level
Johnny Wei and Robin Jia. 2021 · 2021
Cited alongside, same era.
Results of WMT22 metrics shared task: Stop using BLEU – neural metrics are better and more robust
Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, Eleftherios Avramidis, Tom Kocmi, George Foster, Alon Lavie, and André F. T. Martins. 2022 · 2022
Cited alongside, same era.
xcomet: Transparent machine translation evaluation through fine-grained error detection
Nuno M. Guerreiro, Ricardo Rei, Daan van Stigt, Luisa Coheur, Pierre Colombo, and André F. T. Martins. 2023 · 2023
Later among the works it cites.
How good are gpt models at machine translation? a comprehensive evaluation
Amr Hendy, Mohamed Abdelrehim, Amr Sharaf, Vikas Raunak, Mohamed Gabr, Hitokazu Matsushita, Young Jin Kim, Mohamed Afify, and Hany Hassan Awadalla. 2023 · 2023
Later among the works it cites.
MetricX-23: The Google submission to the WMT 2023 metrics shared task
Juraj Juraska, Mara Finkelstein, Daniel Deutsch, Aditya Siddhant, Mehdi Mirzazadeh, and Markus Freitag. 2023 · 2023
Later among the works it cites.
Findings of the 2023 conference on machine translation (WMT23): LLMs are here but not quite there yet
Tom Kocmi, Eleftherios Avramidis, Rachel Bawden, Ondřej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Markus Freitag, Thamme Gowda, Roman Grundkiewicz, Barry Haddow, Philipp Koehn, Benjamin Marie, Christof Monz, Makoto Morishita, Kenton Murray, Makoto Nagata, Toshiaki Nakazawa, Martin Popel, Maja Popović, and Mariya Shmatova. 2023 · 2023
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
On the effectiveness of automated metrics for text generation systems
Pius von Däniken, Jan Deriu, Don Tuggener, and Mark Cieliebak. 2022 · 2022
Cited alongside, same era.
Correction of errors in preference ratings from automated metrics for text generation
Jan Deriu, Pius von Däniken, Don Tuggener, and Mark Cieliebak. 2023 · 2023
Cited alongside, same era.
Results of WMT23 metrics shared task: Metrics might be guilty but references are not innocent
Markus Freitag, Nitika Mathur, Chi-kiu Lo, Eleftherios Avramidis, Ricardo Rei, Brian Thompson, Tom Kocmi, Frederic Blain, Daniel Deutsch, Craig Stewart, Chrysoula Zerva, Sheila Castilho, Alon Lavie, and George Foster. 2023 · 2023
Cited alongside, same era.
Paraphrase generation as zero-shot multilingual translation: Disentangling semantic similarity from lexical and syntactic diversity
Brian Thompson and Matt Post. 2020b
Cited in the paper.
GEMBA-MQM: Detecting translation quality error spans with GPT-4
Tom Kocmi and Christian Federmann. 2023 · 2023
Later among the works it cites.
Exploring prompt engineering with GPT language models for document-level machine translation: Insights and findings
Yangjian Wu and Gang Hu. 2023 · 2023
Later among the works it cites.
Favi-score: A measure for favoritism in automated preference ratings for generative AI evaluation
Pius von Däniken, Jan Deriu, Don Tuggener, and Mark Cieliebak. 2024 · 2024
Closest in time.
Calibrate-extrapolate: Rethinking prevalence estimation with black box classifiers
Siqi Wu and Paul Resnick. 2024 · 2024
Closest in time.