Fetching the paper…
Reading the bibliography…
We present COMET, a neural framework for training multilingual machine translation evaluation models which obtains new state-of-the-art levels of correlation with human judgements.
Machine Translation Evaluation with BERT Regressor
Hiroki Shimanaka, Tomoyuki Kajiwara, and Mamoru Komachi. 2019 · 1907
Earlier work this paper cites.
Distilbert, a distilled version of BERT: smaller, faster, cheaper and lighter
Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. 2019 · 1910
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1911
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
A study of translation edit rate with targeted human annotation
Matthew Snover, Bonnie Dorr, Richard Schwartz, Linnea Micciulla, and John Makhoul. 2006 · 2006
Earlier work this paper cites.
Moses: Open source toolkit for statistical machine translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondřej Bojar, Alexandra Constantin, and Evan Herbst. 2007 · 2007
Earlier work this paper cites.
The meteor metric for automatic evaluation of machine translation
Alon Lavie and Michael Denkowski. 2009 · 2009
Earlier work this paper cites.
Meteor 1.3: Automatic metric for reliable optimization and evaluation of machine translation systems
Michael Denkowski and Alon Lavie. 2011 · 2011
Earlier work this paper cites.
Findings of the 2013 Workshop on Statistical Machine Translation
Ondřej Bojar, Christian Buck, Chris Callison-Burch, Christian Federmann, Barry Haddow, Philipp Koehn, Christof Monz, Matt Post, Radu Soricut, and Lucia Specia. 2013 · 2013
Earlier work this paper cites.
Continuous measurement scales in human evaluation of machine translation
Yvette Graham, Timothy Baldwin, Alistair Moffat, and Justin Zobel. 2013 · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013 · 2013
Earlier work this paper cites.
Findings of the 2014 workshop on statistical machine translation
Ondřej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, Christof Monz, Pavel Pecina, Matt Post, Herve Saint-Amand, Radu Soricut, Lucia Specia, and Aleš Tamchyna. 2014 · 2014
Earlier work this paper cites.
Is machine translation getting better over time?
Yvette Graham, Timothy Baldwin, Alistair Moffat, and Justin Zobel. 2014 · 2014
Earlier work this paper cites.
Multidimensional quality metrics (MQM): A framework for declaring and describing translation quality metrics
Arle Lommel, Aljoscha Burchardt, and Hans Uszkoreit. 2014 · 2014
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Earlier work this paper cites.
Findings of the 2015 workshop on statistical machine translation
Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Barry Haddow, Matthias Huck, Chris Hokamp, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Matt Post, Carolina Scarton, Lucia Specia, and Marco Turchi. 2015 · 2015
Earlier work this paper cites.
chrF: character n-gram f-score for automatic MT evaluation
Maja Popović. 2015 · 2015
Earlier work this paper cites.
Facenet: A unified embedding for face recognition and clustering
F. Schroff, D. Kalenichenko, and J. Philbin. 2015 · 2015
Cited alongside, same era.
Findings of the 2016 conference on machine translation
Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Aurélie Névéol, Mariana Neves, Martin Popel, Matt Post, Raphael Rubino, Carolina Scarton, Lucia Specia, Marco Turchi, Karin Verspoor, and Marcos Zampieri. 2016 · 2016
Cited alongside, same era.
Word2Vec vs DBnary: Augmenting METEOR using vector representations or lexical resources?
Christophe Servan, Alexandre Bérard, Zied Elloumi, Hervé Blanchon, and Laurent Besacier. 2016 · 2016
Cited alongside, same era.
Findings of the 2017 conference on machine translation (WMT17)
Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Shujian Huang, Matthias Huck, Philipp Koehn, Qun Liu, Varvara Logacheva, Christof Monz, Matteo Negri, Matt Post, Raphael Rubino, Lucia Specia, and Marco Turchi. 2017a · 2017
Cited alongside, same era.
Can machine translation systems be evaluated by the crowd alone
Yvette Graham, Timothy Baldwin, Alistair Moffat, and Justin Zobel. 2017 · 2017
PyTorch Lightning: The lightweight PyTorch wrapper for high-performance AI research
WA Falcon. 2019 · 2019
Later among the works it cites.
Findings of the WMT 2019 shared tasks on quality estimation
Erick Fonseca, Lisa Yankovskaya, André F. T. Martins, Mark Fishel, and Christian Federmann. 2019 · 2019
Later among the works it cites.
Meteor++ 2.0: Adopt syntactic level paraphrase knowledge into machine translation evaluation
Yinuo Guo and Junfeng Hu. 2019 · 2019
Later among the works it cites.
75 languages, 1 model: Parsing universal dependencies universally
Dan Kondratyuk and Milan Straka. 2019 · 2019
Later among the works it cites.
YiSi - a unified semantic MT quality evaluation and estimation metric for languages with different levels of available resources
Chi-kiu Lo. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
chrF++: words helping character n-grams
Maja Popović. 2017 · 2017
Cited alongside, same era.
Translation quality and productivity: A study on rich morphology languages
Lucia Specia, Kim Harris, Frédéric Blain, Aljoscha Burchardt, Viviven Macketanz, Inguna Skadina, Matteo Negri, , and Marco Turchi. 2017 · 2017
Cited alongside, same era.
bleu2vec: the painfully familiar metric on continuous vector space steroids
Andre Tättar and Mark Fishel. 2017 · 2017
Cited alongside, same era.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder. 2018 · 2018
Cited alongside, same era.
Results of the WMT18 metrics shared task: Both characters and embeddings achieve good performance
Qingsong Ma, Ondřej Bojar, and Yvette Graham. 2018 · 2018
Cited alongside, same era.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Cited alongside, same era.
RUSE: Regressor using sentence embeddings for automatic machine translation evaluation
Hiroki Shimanaka, Tomoyuki Kajiwara, and Mamoru Komachi. 2018 · 2018
Cited alongside, same era.
Qingsong Ma, Johnny Wei, Ondřej Bojar, and Yvette Graham. 2019 · 2019
Later among the works it cites.
Putting evaluation in context: Contextual embeddings improve machine translation evaluation
Nitika Mathur, Timothy Baldwin, and Trevor Cohn. 2019 · 2019
Later among the works it cites.
How multilingual is multilingual BERT?
Telmo Pires, Eva Schlinger, and Dan Garrette. 2019 · 2019
Later among the works it cites.
Sentence-BERT: Sentence embeddings using Siamese BERT-networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Later among the works it cites.
BERT rediscovers the classical NLP pipeline
Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019 · 2019
Later among the works it cites.
MoverScore: Text generation evaluating with contextualized embeddings and earth mover distance
Wei Zhao, Maxime Peyrard, Fei Liu, Yang Gao, Christian M. Meyer, and Steffen Eger. 2019 · 2019
Later among the works it cites.
Practical Guidelines for the Use of MQM in Scientific Research on Translation quality
Aljoscha Burchardt and Arle Lommel. 2014 · 2020
Closest in time.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Closest in time.
Automatic machine translation evaluation using source language inputs and cross-lingual language model
Kosuke Takahashi, Katsuhito Sudoh, and Satoshi Nakamura. 2020 · 2020
Closest in time.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Closest in time.
On the limitations of cross-lingual encoders as exposed by reference-free machine translation evaluation
Wei Zhao, Goran Glavaš, Maxime Peyrard, Yang Gao, Robert West, and Steffen Eger. 2020 · 2020
Closest in time.