Fetching the paper…
Reading the bibliography…
Widely used learned metrics for machine translation evaluation, such as COMET and BLEURT, estimate the quality of a translation hypothesis by providing a single sentence-level score.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
A study of translation edit rate with targeted human annotation
Matthew Snover, Bonnie Dorr, Richard Schwartz, Linnea Micciulla, and John Makhoul. 2006 · 2006
Earlier work this paper cites.
Machine translation evaluation versus quality estimation
Lucia Specia, Dhwaj Raj, and Marco Turchi. 2010 · 2010
Earlier work this paper cites.
Continuous measurement scales in human evaluation of machine translation
Yvette Graham, Timothy Baldwin, Alistair Moffat, and Justin Zobel. 2013 · 2013
Earlier work this paper cites.
Multidimensional quality metrics (MQM): A framework for declaring and describing translation quality metrics
Arle Lommel, Aljoscha Burchardt, and Hans Uszkoreit. 2014 · 2014
Earlier work this paper cites.
chrF: character n-gram F-score for automatic MT evaluation
Maja Popović. 2015 · 2015
Earlier work this paper cites.
Findings of the 2017 conference on machine translation (WMT17)
Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Shujian Huang, Matthias Huck, Philipp Koehn, Qun Liu, Varvara Logacheva, Christof Monz, Matteo Negri, Matt Post, Raphael Rubino, Lucia Specia, and Marco Turchi. 2017 · 2017
Earlier work this paper cites.
Optuna: A next-generation hyperparameter optimization framework
Takuya Akiba, Shotaro Sano, Toshihiko Yanase, Takeru Ohta, and Masanori Koyama. 2019 · 2019
Earlier work this paper cites.
Findings of the WMT 2019 shared tasks on quality estimation
Erick Fonseca, Lisa Yankovskaya, André F. T. Martins, Mark Fishel, and Christian Federmann. 2019 · 2019
Earlier work this paper cites.
COMET: A neural framework for MT evaluation
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020 · 2020
Earlier work this paper cites.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Earlier work this paper cites.
A statistical analysis of summarization evaluation metrics using resampling methods
Daniel Deutsch, Rotem Dror, and Dan Roth. 2021 · 2021
Earlier work this paper cites.
Beyond english-centric multilingual machine translation
Angela Fan, Shruti Bhosale, Holger Schwenk, Zhiyi Ma, Ahmed El-Kishky, Siddharth Goyal, Mandeep Baines, Onur Celebi, Guillaume Wenzek, Vishrav Chaudhary, Naman Goyal, Tom Birch, Vitaliy Liptchinsky, Sergey Edunov, Edouard Grave, Michael Auli, and Armand Joulin. 2021 · 2021
Earlier work this paper cites.
Larger-scale transformers for multilingual masked language modeling
Naman Goyal, Jingfei Du, Myle Ott, Giri Anantharaman, and Alexis Conneau. 2021 · 2021
Earlier work this paper cites.
To ship or not to ship: An extensive evaluation of automatic metrics for machine translation
Tom Kocmi, Christian Federmann, Roman Grundkiewicz, Marcin Junczys-Dowmunt, Hitokazu Matsushita, and Arul Menezes. 2021 · 2021
Earlier work this paper cites.
Scientific credibility of machine translation research: A meta-evaluation of 769 papers
Benjamin Marie, Atsushi Fujita, and Raphael Rubino. 2021 · 2021
Cited alongside, same era.
Learning compact metrics for MT
Amy Pu, Hyung Won Chung, Ankur Parikh, Sebastian Gehrmann, and Thibault Sellam. 2021 · 2021
Cited alongside, same era.
The curious case of hallucinations in neural machine translation
Vikas Raunak, Arul Menezes, and Marcin Junczys-Dowmunt. 2021 · 2021
Cited alongside, same era.
mT5: A massively multilingual pre-trained text-to-text transformer
Linting Xue, Noah Constant, Adam Roberts, Mihir Kale, Rami Al-Rfou, Aditya Siddhant, Aditya Barua, and Colin Raffel. 2021 · 2021
Cited alongside, same era.
Robust MT evaluation with sentence-level multilingual augmentation
Duarte Alves, Ricardo Rei, Ana C Farinha, José G. C. de Souza, and André F. T. Martins. 2022 · 2022
Cited alongside, same era.
A unified dialogue user simulator for few-shot data augmentation
Dazhen Wan, Zheng Zhang, Qi Zhu, Lizi Liao, and Minlie Huang. 2022a · 2022
Later among the works it cites.
Findings of the WMT 2022 shared task on quality estimation
Chrysoula Zerva, Frédéric Blain, Ricardo Rei, Piyawat Lertvittayakumjorn, José G. C. de Souza, Steffen Eger, Diptesh Kanojia, Duarte Alves, Constantin Orăsan, Marina Fomicheva, André F. T. Martins, and Lucia Specia. 2022 · 2022
Later among the works it cites.
Towards fine-grained information: Identifying the type and location of translation errors
Keqin Bao, Yu Wan, Dayiheng Liu, Baosong Yang, Wenqiang Lei, Xiangnan He, Derek F. Wong, and Jun Xie. 2023 · 2023
Closest in time.
David Dale, Elena Voita, Janice Lam, Prangthip Hansanti, Christophe Ropers, Elahe Kalbassi, Cynthia Gao, Loïc Barrault, and Marta R. Costa-jussà. 2023 · 2023
Closest in time.
Training and meta-evaluating machine translation evaluation metrics at the paragraph level
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chantal Amrhein and Rico Sennrich. 2022 · 2022
Cited alongside, same era.
Language-agnostic BERT sentence embedding
Fangxiaoyu Feng, Yinfei Yang, Daniel Cer, Naveen Arivazhagan, and Wei Wang. 2022 · 2022
Cited alongside, same era.
Quality-aware decoding for neural machine translation
Patrick Fernandes, António Farinhas, Ricardo Rei, José G. C. de Souza, Perez Ogayo, Graham Neubig, and Andre Martins. 2022 · 2022
Cited alongside, same era.
MLQE-PE: A multilingual quality estimation and post-editing dataset
Marina Fomicheva, Shuo Sun, Erick Fonseca, Chrysoula Zerva, Frédéric Blain, Vishrav Chaudhary, Francisco Guzmán, Nina Lopatina, Lucia Specia, and André F. T. Martins. 2022 · 2022
Cited alongside, same era.
Results of WMT22 metrics shared task: Stop using BLEU – neural metrics are better and more robust
Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, Eleftherios Avramidis, Tom Kocmi, George Foster, Alon Lavie, and André F. T. Martins. 2022 · 2022
Cited alongside, same era.
DEMETR: Diagnosing evaluation metrics for translation
Marzena Karpinska, Nishant Raj, Katherine Thai, Yixiao Song, Ankita Gupta, and Mohit Iyyer. 2022 · 2022
Cited alongside, same era.
MaTESe: Machine translation evaluation as a sequence tagging problem
Stefano Perrella, Lorenzo Proietti, Alessandro Scirè, Niccolò Campolungo, and Roberto Navigli. 2022 · 2022
Cited alongside, same era.
Daniel Deutsch, Juraj Juraska, Mara Finkelstein, and Markus Freitag. 2023 · 2023
Closest in time.
Patrick Fernandes, Daniel Deutsch, Mara Finkelstein, Parker Riley, André F. T. Martins, Graham Neubig, Ankush Garg, Jonathan H. Clark, Markus Freitag, and Orhan Firat. 2023 · 2023
Closest in time.
Reinforced self-training (rest) for language modeling
Caglar Gulcehre, Tom Le Paine, Srivatsan Srinivasan, Ksenia Konyushkova, Lotte Weerts, Abhishek Sharma, Aditya Siddhant, Alex Ahern, Miaosen Wang, Chenjie Gu, Wolfgang Macherey, Arnaud Doucet, Orhan Firat, and Nando de Freitas. 2023 · 2023
Closest in time.
Large language models are state-of-the-art evaluators of translation quality
Tom Kocmi and Christian Federmann. 2023 · 2023
Closest in time.
Towards explainable evaluation metrics for machine translation
Christoph Leiter, Piyawat Lertvittayakumjorn, M. Fomicheva, Wei Zhao, Yang Gao, and Steffen Eger. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.
Scaling up cometkiwi: Unbabel-ist 2023 submission for the quality estimation shared task
Ricardo Rei, Nuno M. Guerreiro, José Pombal, Daan van Stigt, Marcos Treviso, Luisa Coheur, José G. C. de Souza, and André F. T. Martins. 2023a · 2023
Closest in time.
IndicMT eval: A dataset to meta-evaluate machine translation metrics for Indian languages
Ananya Sai B, Tanay Dixit, Vignesh Nagarajan, Anoop Kunchukuttan, Pratyush Kumar, Mitesh M. Khapra, and Raj Dabre. 2023 · 2023
Closest in time.
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timothée Lacroix, Baptiste Rozière, Naman Goyal, Eric Hambro, Faisal Azhar, Aurelien Rodriguez, Armand Joulin, Edouard Grave, and Guillaume Lample. 2023 · 2023
Closest in time.
Instructscore: Towards explainable text generation evaluation with automatic feedback
Wenda Xu, Danqing Wang, Liangming Pan, Zhenqiao Song, Markus Freitag, William Yang Wang, and Lei Li. 2023 · 2023
Closest in time.