Fetching the paper…
Reading the bibliography…
Automatic machine translation (MT) metrics are widely used to distinguish the translation qualities of machine translation systems across relatively large test sets (system-level evaluation).
RoBERTa: A Robustly Optimized BERT Pretraining Approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Comparison of the predicted and observed secondary structure of t4 phage lysozyme
Brian W. Matthews. 1975 · 1975
Earlier work this paper cites.
The ATIS spoken language systems pilot corpus
Charles T. Hemphill, John J. Godfrey, and George R. Doddington. 1990 · 1990
Earlier work this paper cites.
The ARPA MT evaluation methodologies: Evolution, lessons, and future approaches
John S. White, Theresa A. O’Connell, and Francis E. O’Mara. 1994 · 1994
Earlier work this paper cites.
Evaluating Natural Language Processing Systems, An Analysis and Review , volume 1083 of Lecture Notes in Computer Science
Karen Sparck Jones and Julia Rose Galliers, editors. 1996 · 1996
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
The Error Is the Clue: Breakdown In Human-Machine Interaction
Bilyana Martinovski and David Traum. 2003 · 2003
Earlier work this paper cites.
Re-evaluating the role of Bleu in machine translation research
Chris Callison-Burch, Miles Osborne, and Philipp Koehn. 2006 · 2006
Earlier work this paper cites.
Manual and automatic evaluation of machine translation between European languages
Philipp Koehn and Christof Monz. 2006 · 2006
Earlier work this paper cites.
Task-based MT evaluation: From who/when/where extraction to event understanding
Jamal Laoudi, Calandra R. Tate, and Clare R. Voss. 2006 · 2006
Earlier work this paper cites.
(meta-) evaluation of machine translation
Chris Callison-Burch, Cameron Fordyce, Philipp Koehn, Christof Monz, and Josh Schroeder. 2007 · 2007
Earlier work this paper cites.
Human evaluation of machine translation through binary system comparisons
David Vilar, Gregor Leusch, Hermann Ney, and Rafael E. Banchs. 2007 · 2007
Earlier work this paper cites.
Continuous measurement scales in human evaluation of machine translation
Yvette Graham, Timothy Baldwin, Alistair Moffat, and Justin Zobel. 2013 · 2013
Earlier work this paper cites.
Findings of the 2014 workshop on statistical machine translation
Ondřej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, Christof Monz, Pavel Pecina, Matt Post, Herve Saint-Amand, Radu Soricut, Lucia Specia, and Aleš Tamchyna. 2014 · 2014
Earlier work this paper cites.
Multidimensional quality metrics (mqm): A framework for declaring and describing translation quality metrics
Arle Lommel, Aljoscha Burchardt, and Hans Uszkoreit. 2014 · 2014
Earlier work this paper cites.
Results of the WMT15 metrics shared task
Miloš Stanojević, Amir Kamran, Philipp Koehn, and Ondřej Bojar. 2015 · 2015
Earlier work this paper cites.
SQuAD: 100,000+ questions for machine comprehension of text
Pranav Rajpurkar, Jian Zhang, Konstantin Lopyrev, and Percy Liang. 2016 · 2016
Earlier work this paper cites.
Results of the WMT17 metrics shared task
Ondřej Bojar, Yvette Graham, and Amir Kamran. 2017 · 2017
Earlier work this paper cites.
Learning a neural semantic parser from user feedback
Srinivasan Iyer, Ioannis Konstas, Alvin Cheung, Jayant Krishnamurthy, and Luke Zettlemoyer. 2017 · 2017
Earlier work this paper cites.
chrF++: words helping character n-grams
Maja Popović. 2017 · 2017
Earlier work this paper cites.
Automatic reference-based evaluation of pronoun translation misses the point
Liane Guillou and Christian Hardmeier. 2018 · 2018
Cited alongside, same era.
SUMBT: Slot-utterance matching for universal and scalable belief tracking
Hwaran Lee, Jinsik Lee, and Tae-Yoon Kim. 2019 · 2019
Cited alongside, same era.
Results of the WMT19 metrics shared task: Segment-level and strong MT systems pose big challenges
Qingsong Ma, Johnny Wei, Ondřej Bojar, and Yvette Graham. 2019 · 2019
Cited alongside, same era.
Estimating post-editing effort: a study on human judgements, task-based and reference-based metrics of MT quality
Scarton Scarton, Mikel L. Forcada, Miquel Esplà-Gomis, and Lucia Specia. 2019 · 2019
Cited alongside, same era.
On the cross-lingual transferability of monolingual representations
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2020 · 2020
Cited alongside, same era.
To ship or not to ship: An extensive evaluation of automatic metrics for machine translation
Tom Kocmi, Christian Federmann, Roman Grundkiewicz, Marcin Junczys-Dowmunt, Hitokazu Matsushita, and Arul Menezes. 2021 · 2021
Later among the works it cites.
Agree to disagree: Analysis of inter-annotator disagreements in human evaluation of machine translation output
Maja Popović. 2021 · 2021
Later among the works it cites.
XTREME-R: Towards more challenging and nuanced multilingual evaluation
Sebastian Ruder, Noah Constant, Jan Botha, Aditya Siddhant, Orhan Firat, Jinlan Fu, Pengfei Liu, Junjie Hu, Dan Garrette, Graham Neubig, and Melvin Johnson. 2021 · 2021
Later among the works it cites.
Multilingual translation from denoising pre-training
Yuqing Tang, Chau Tran, Xian Li, Peng-Jen Chen, Naman Goyal, Vishrav Chaudhary, Jiatao Gu, and Angela Fan. 2021 · 2021
Later among the works it cites.
Neural machine translation quality and post-editing performance
Vilém Zouhar, Martin Popel, Ondřej Bojar, and Aleš Tamchyna. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
A study in improving BLEU reference coverage with diverse automatic paraphrasing
Rachel Bawden, Biao Zhang, Lisa Yankovskaya, Andre Tättar, and Matt Post. 2020 · 2020
Cited alongside, same era.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Cited alongside, same era.
MultiWOZ 2.1: A consolidated multi-domain dialogue dataset with state corrections and state tracking baselines
Mihail Eric, Rahul Goel, Shachi Paul, Abhishek Sethi, Sanchit Agarwal, Shuyang Gao, Adarsh Kumar, Anuj Goyal, Peter Ku, and Dilek Hakkani-Tur. 2020 · 2020
Cited alongside, same era.
BLEU might be guilty but references are not innocent
Markus Freitag, David Grangier, and Isaac Caswell. 2020 · 2020
Cited alongside, same era.
Statistical power and translationese in machine translation evaluation
Yvette Graham, Barry Haddow, and Philipp Koehn. 2020 · 2020
Cited alongside, same era.
XTREME: A massively multilingual multi-task benchmark for evaluating cross-lingual generalisation
Junjie Hu, Sebastian Ruder, Aditya Siddhant, Graham Neubig, Orhan Firat, and Melvin Johnson. 2020 · 2020
Cited alongside, same era.
COMET: A neural framework for MT evaluation
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020 · 2020
Cited alongside, same era.
ACES: Translation Accuracy Challenge Sets for Evaluating Machine Translation Metrics
Chantal Amrhein, Nikita Moghe, and Liane Guillou. 2022 · 2022
Closest in time.
Identifying weaknesses in machine translation metrics through minimum Bayes risk decoding: A case study for COMET
Chantal Amrhein and Rico Sennrich. 2022 · 2022
Closest in time.
Indicxtreme: A multi-task benchmark for evaluating indic languages
Sumanth Doddapaneni, Rahul Aralikatte, Gowtham Ramesh, Shreya Goyal, Mitesh M. Khapra, Anoop Kunchukuttan, and Pratyush Kumar. 2022 · 2022
Closest in time.
MLQE-PE: A multilingual quality estimation and post-editing dataset
Marina Fomicheva, Shuo Sun, Erick Fonseca, Chrysoula Zerva, Frédéric Blain, Vishrav Chaudhary, Francisco Guzmán, Nina Lopatina, Lucia Specia, and André F. T. Martins. 2022 · 2022
Closest in time.
Multi2WOZ: A robust multilingual dataset and conversational pretraining for task-oriented dialog
Chia-Chien Hung, Anne Lauscher, Ivan Vulić, Simone Ponzetto, and Goran Glavaš. 2022 · 2022
Closest in time.
An Automatic Evaluation of the WMT22 General Machine Translation Task
Benjamin Marie. 2022 · 2022
Closest in time.
Overview of the 9th workshop on Asian translation
Toshiaki Nakazawa, Hideya Mino, Isao Goto, Raj Dabre, Shohei Higashiyama, Shantipriya Parida, Anoop Kunchukuttan, Makoto Morishita, Ondřej Bojar, Chenhui Chu, Akiko Eriguchi, Kaori Abe, Yusuke Oda, and Sadao Kurohashi. 2022 · 2022
Closest in time.
MaTESe: Machine Translation Evaluation as a Sequence Tagging Problem
Stefano Perrella, Lorenzo Proietti, Alessandro Scirè, Niccolò Campolungo, and Roberto Navigli. 2022 · 2022
Closest in time.
COMET-22: Unbabel-IST 2022 Submission for the Metrics Shared Task
Ricardo Rei, José G. C. de Souza, Duarte Alves, Chrysoula Zerva, Ana C Farinha, Taisiya Glushkova, Alon Lavie, Luisa Coheur, and André F. T. Martins. 2022 · 2022
Closest in time.
Zero-shot cross-lingual semantic parsing
Tom Sherborne and Mirella Lapata. 2022 · 2022
Closest in time.
UniTE: Unified translation evaluation
Yu Wan, Dayiheng Liu, Baosong Yang, Haibo Zhang, Boxing Chen, Derek Wong, and Lidia Chao. 2022 · 2022
Closest in time.
Evaluating machine translation in cross-lingual E-commerce search
Hang Zhang, Liling Tan, and Amita Misra. 2022 · 2022
Closest in time.
Cross-Lingual Dialogue Dataset Creation via Outline-Based Generation
Olga Majewska, Evgeniia Razumovskaia, Edoardo M. Ponti, Ivan Vulić, and Anna Korhonen. 2023 · 2023
Closest in time.
Meta-Learning a Cross-lingual Manifold for Semantic Parsing
Tom Sherborne and Mirella Lapata. 2023 · 2023
Closest in time.