Fetching the paper…
Reading the bibliography…
Annually, research teams spend large amounts of money to evaluate the quality of machine translation systems (WMT, inter alia).
Multidimensional quality metrics (MQM): A framework for declaring and describing translation quality metrics
Arle Lommel, Hans Uszkoreit, and Aljoscha Burchardt. 2014 · 2014
Earlier work this paper cites.
Accurate evaluation of segment-level machine translation metrics
Yvette Graham, Timothy Baldwin, and Nitika Mathur. 2015 · 2015
Earlier work this paper cites.
Appraise evaluation framework for machine translation
Christian Federmann. 2018 · 2018
Earlier work this paper cites.
RankME: Reliable human ratings for natural language generation
Jekaterina Novikova, Ondřej Dušek, and Verena Rieser. 2018 · 2018
Earlier work this paper cites.
BLEU might be guilty but references are not innocent
Markus Freitag, David Grangier, and Isaac Caswell. 2020 · 2020
Earlier work this paper cites.
To trust or to think: Cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making
Zana Buçinca, Maja Barbara Malaya, and Krzysztof Z Gajos. 2021 · 2021
Earlier work this paper cites.
The Eval4NLP shared task on explainable quality estimation: Overview and results
Marina Fomicheva, Piyawat Lertvittayakumjorn, Wei Zhao, Steffen Eger, and Yang Gao. 2021 · 2021
Earlier work this paper cites.
Experts, errors, and context: A large-scale study of human evaluation for machine translation
Markus Freitag, George Foster, David Grangier, Viresh Ratnakar, Qijun Tan, and Wolfgang Macherey. 2021 · 2021
Earlier work this paper cites.
Neural machine translation quality and post-editing performance
Vilém Zouhar, Martin Popel, Ondřej Bojar, and Aleš Tamchyna. 2021 · 2021
Earlier work this paper cites.
Role of human-AI interaction in selective prediction
Elizabeth Bondi, Raphael Koster, Hannah Sheahan, Martin Chadwick, Yoram Bachrach, Taylan Cemgil, Ulrich Paquet, and Krishnamurthy Dvijotham. 2022 · 2022
Cited alongside, same era.
Design-for-responsible algorithmic decision-making systems: A question of ethical judgement and human meaningful control
W David Holford. 2022 · 2022
Cited alongside, same era.
Findings of the 2022 conference on machine translation (WMT22)
Tom Kocmi, Rachel Bawden, Ondřej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Thamme Gowda, Yvette Graham, Roman Grundkiewicz, Barry Haddow, Rebecca Knowles, Philipp Koehn, Christof Monz, Makoto Morishita, Masaaki Nagata, Toshiaki Nakazawa, Michal Novák, Martin Popel, and Maja Popović. 2022 · 2022
Cited alongside, same era.
Taglab: AI-assisted annotation for the fast and accurate semantic segmentation of coral reef orthoimages
Gaia Pavoni, Massimiliano Corsini, Federico Ponchio, Alessandro Muntoni, Clinton Edwards, Nicole Pedersen, Stuart Sandin, and Paolo Cignoni. 2022 · 2022
Cited alongside, same era.
AI-Assisted deep NLP-Based approach for prediction of fake news from social media users
Findings of the 2023 conference on machine translation (WMT23): LLMs are here but not quite there yet
Tom Kocmi, Eleftherios Avramidis, Rachel Bawden, Ondřej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Markus Freitag, Thamme Gowda, Roman Grundkiewicz, Barry Haddow, Philipp Koehn, Benjamin Marie, Christof Monz, Makoto Morishita, Kenton Murray, Makoto Nagata, Toshiaki Nakazawa, Martin Popel, Maja Popović, and Mariya Shmatova. 2023 · 2023
Later among the works it cites.
GEMBA-MQM: Detecting translation quality error spans with GPT-4
Tom Kocmi and Christian Federmann. 2023 · 2023
Later among the works it cites.
G-eval: NLG evaluation using gpt-4 with better human alignment
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. 2023 · 2023
Later among the works it cites.
Veniamin Veselovsky, Manoel Horta Ribeiro, and Robert West. 2023 · 2023
Later among the works it cites.
COMET for low-resource machine translation evaluation: A case study of English-Maltese and Spanish-Basque
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ganesh Gopal Devarajan, Senthil Murugan Nagarajan, Sardar Irfanullah Amanullah, SA Sahaaya Arul Mary, and Ali Kashif Bashir. 2023 · 2023
Cited alongside, same era.
A diachronic perspective on user trust in AI under uncertainty
Shehzaad Dhuliawala, Vilém Zouhar, Mennatallah El-Assady, and Mrinmaya Sachan. 2023 · 2023
Cited alongside, same era.
The devil is in the errors: Leveraging large language models for fine-grained machine translation evaluation
Patrick Fernandes, Daniel Deutsch, Mara Finkelstein, Parker Riley, André Martins, Graham Neubig, Ankush Garg, Jonathan Clark, Markus Freitag, and Orhan Firat. 2023 · 2023
Cited alongside, same era.
Results of WMT23 metrics shared task: Metrics might be guilty but references are not innocent
Markus Freitag, Nitika Mathur, Chi-kiu Lo, Eleftherios Avramidis, Ricardo Rei, Brian Thompson, Tom Kocmi, Frederic Blain, Daniel Deutsch, Craig Stewart, Chrysoula Zerva, Sheila Castilho, Alon Lavie, and George Foster. 2023 · 2023
Cited alongside, same era.
xCOMET: Transparent machine translation evaluation through fine-grained error detection
Nuno M. Guerreiro, Ricardo Rei, Daan van Stigt, Luisa Coheur, Pierre Colombo, and André F. T. Martins. 2023 · 2023
Cited alongside, same era.
Findings of the WMT24 general machine translation shared task: The LLM era is here but MT is not solved yet
Tom Kocmi, Eleftherios Avramidis, Rachel Bawden, Ondřej Bojar, Anton Dvorkovich, Christian Federmann, Mark Fishel, Markus Freitag, Thamme Gowda, Roman Grundkiewicz, Barry Haddow, Marzena Karpinska, Philipp Koehn, Benjamin Marie, Christof Monz, Kenton Murray, Masaaki Nagata, Martin Popel, Maja Popović, Mariya Shmatova, Steinthór Steingrímsson, and Vilém Zouhar. 2024a
Cited in the paper.
Error span annotation: A balanced approach for human evaluation of machine translation
Tom Kocmi, Vilém Zouhar, Eleftherios Avramidis, Roman Grundkiewicz, Marzena Karpinska, Maja Popović, Mrinmaya Sachan, and Mariya Shmatova. 2024b
Cited in the paper.
Júlia Falcão, Claudia Borg, Nora Aranberri, and Kurt Abela. 2024 · 2024
Closest in time.
Finding replicable human evaluations via stable ranking probability
Parker Riley, Daniel Deutsch, George Foster, Viresh Ratnakar, Ali Dabirmoghaddam, and Markus Freitag. 2024 · 2024
Closest in time.
Quality and quantity of machine translation references for automatic metrics
Vilém Zouhar and Ondřej Bojar. 2024 · 2024
Closest in time.
Fine-tuned machine translation metrics struggle in unseen domains
Vilém Zouhar, Shuoyang Ding, Anna Currey, Tatyana Badeka, Jenyuan Wang, and Brian Thompson. 2024 · 2024
Closest in time.