Fetching the paper…
Reading the bibliography…
We frame the task of machine translation evaluation as one of scoring machine translation output with a sequence-to-sequence paraphraser, conditioned on a human reference.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019 · 1904
Earlier work this paper cites.
Massively multilingual neural machine translation in the wild: Findings and challenges
Naveen Arivazhagan, Ankur Bapna, Orhan Firat, Dmitry Lepikhin, Melvin Johnson, Maxim Krikun, Mia Xu Chen, Yuan Cao, George Foster, Colin Cherry, Wolfgang Macherey, Zhifeng Chen, and Yonghui Wu. 2019 · 1907
Earlier work this paper cites.
WikiMatrix: Mining 135m parallel sentences in 1620 language pairs from wikipedia
Holger Schwenk, Vishrav Chaudhary, Shuo Sun, Hongyu Gong, and Francisco Guzmán. 2019 · 1907
Earlier work this paper cites.
Automatic evaluation of machine translation quality using n-gram co-occurrence statistics
George Doddington. 2002 · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Statistical significance tests for machine translation evaluation
Philipp Koehn. 2004 · 2004
Earlier work this paper cites.
Monolingual machine translation for paraphrase generation
Chris Quirk, Chris Brockett, and William Dolan. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Paraphrasing with bilingual parallel corpora
Colin Bannard and Chris Callison-Burch. 2005 · 2005
Earlier work this paper cites.
Europarl: A parallel corpus for statistical machine translation
Philipp Koehn. 2005 · 2005
Earlier work this paper cites.
CDER: Efficient MT evaluation using block movements
Gregor Leusch, Nicola Ueffing, and Hermann Ney. 2006 · 2006
Earlier work this paper cites.
A study of translation edit rate with targeted human annotation
Matthew Snover, Bonnie Dorr, Richard Schwartz, Linnea Micciulla, and John Makhoul. 2006 · 2006
Earlier work this paper cites.
ParaEval: Using paraphrases to evaluate summaries automatically
Liang Zhou, Chin-Yew Lin, Dragos Stefan Munteanu, and Eduard Hovy. 2006 · 2006
Earlier work this paper cites.
Robust machine translation evaluation with entailment features
Sebastian Padó, Michel Galley, Dan Jurafsky, and Christopher D. Manning. 2009 · 2009
Earlier work this paper cites.
Extending the METEOR machine translation evaluation metric to the phrase level
Michael Denkowski and Alon Lavie. 2010 · 2010
Earlier work this paper cites.
Multiun: A multilingual corpus from united nation documents
Andreas Eisele and Yu Chen. 2010 · 2010
Earlier work this paper cites.
Simple English Wikipedia: A new text simplification task
William Coster and David Kauchak. 2011 · 2011
Earlier work this paper cites.
MEANT: An inexpensive, high-accuracy, semi-automatic metric for evaluating translation utility based on semantic roles
Chi-kiu Lo and Dekai Wu. 2011 · 2011
Earlier work this paper cites.
Evaluation without references: IBM1 scores as evaluation metrics
Maja Popović, David Vilar, Eleftherios Avramidis, and Aljoscha Burchardt. 2011 · 2011
Earlier work this paper cites.
HyTER: Meaning-equivalent semantics for translation evaluation
Markus Dreyer and Daniel Marcu. 2012 · 2012
Earlier work this paper cites.
LEPOR: A robust evaluation metric for machine translation with augmented factors
Aaron L. F. Han, Derek F. Wong, and Lidia S. Chao. 2012 · 2012
Earlier work this paper cites.
Squibs: What is a paraphrase?
Rahul Bhagat and Eduard Hovy. 2013 · 2013
Earlier work this paper cites.
Paraphrase-driven learning for open question answering
Anthony Fader, Luke Zettlemoyer, and Oren Etzioni. 2013 · 2013
Earlier work this paper cites.
PPDB: The paraphrase database
Juri Ganitkevitch, Benjamin Van Durme, and Chris Callison-Burch. 2013 · 2013
Earlier work this paper cites.
Mt summit13.language-independent model for machine translation evaluation with reinforced factors
Aaron L.-F Han, Derek Wong, Lidia Chao, Liangye He, Yi Lu, Junwen Xing, and Xiaodong Zeng. 2013 · 2013
Earlier work this paper cites.
The multilingual paraphrase database
Juri Ganitkevitch and Chris Callison-Burch. 2014 · 2014
Earlier work this paper cites.
Randomized significance tests in machine translation
Yvette Graham, Nitika Mathur, and Timothy Baldwin. 2014 · 2014
Earlier work this paper cites.
Microsoft COCO: Common objects in context
Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. 2014 · 2014
Earlier work this paper cites.
Multi-task learning for multiple language translation
Daxiang Dong, Hua Wu, Wei He, Dianhai Yu, and Haifeng Wang. 2015 · 2015
Earlier work this paper cites.
ReVal: A simple and effective machine translation evaluation metric based on recurrent neural networks
Rohit Gupta, Constantin Orăsan, and Josef van Genabith. 2015 · 2015
Earlier work this paper cites.
chrF: character n-gram f-score for automatic MT evaluation
Maja Popović. 2015 · 2015
Earlier work this paper cites.
BEER 1.1: ILLC UvA submission to metrics and tuning task
Miloš Stanojević and Khalil Sima’an. 2015 · 2015
Cited alongside, same era.
Reference bias in monolingual machine translation evaluation
Marina Fomicheva and Lucia Specia. 2016 · 2016
Cited alongside, same era.
Bag of tricks for efficient text classification
Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. 2016 · 2016
Cited alongside, same era.
Neural paraphrase generation with stacked residual LSTM networks
Aaditya Prakash, Sadid A. Hasan, Kathy Lee, Vivek Datla, Ashequl Qadir, Joey Liu, and Oladimeji Farri. 2016 · 2016
Cited alongside, same era.
CharacTer: Translation edit rate on character level
Weiyue Wang, Jan-Thorsten Peter, Hendrik Rosendahl, and Hermann Ney. 2016 · 2016
Cited alongside, same era.
Transfer learning for low-resource neural machine translation
RUSE: Regressor using sentence embeddings for automatic machine translation evaluation
Hiroki Shimanaka, Tomoyuki Kajiwara, and Mamoru Komachi. 2018 · 2018
Later among the works it cites.
ParaNMT-50M: Pushing the limits of paraphrastic sentence embeddings with millions of machine translations
John Wieting and Kevin Gimpel. 2018 · 2018
Later among the works it cites.
Massively multilingual neural machine translation
Roee Aharoni, Melvin Johnson, and Orhan Firat. 2019 · 2019
Later among the works it cites.
Low-resource corpus filtering using multilingual sentence embeddings
Vishrav Chaudhary, Yuqing Tang, Francisco Guzmán, Holger Schwenk, and Philipp Koehn. 2019 · 2019
Later among the works it cites.
WMDO: Fluency-based word mover’s distance for machine translation evaluation
Julian Chow, Lucia Specia, and Pranava Madhyastha. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Barret Zoph, Deniz Yuret, Jonathan May, and Kevin Knight. 2016 · 2016
Cited alongside, same era.
Findings of the 2017 conference on machine translation (WMT17)
Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Shujian Huang, Matthias Huck, Philipp Koehn, Qun Liu, Varvara Logacheva, Christof Monz, Matteo Negri, Matt Post, Raphael Rubino, Lucia Specia, and Marco Turchi. 2017 · 2017
Cited alongside, same era.
Enhanced LSTM for natural language inference
Qian Chen, Xiaodan Zhu, Zhen-Hua Ling, Si Wei, Hui Jiang, and Diana Inkpen. 2017 · 2017
Cited alongside, same era.
UHH submission to the WMT17 metrics shared task
Melania Duma and Wolfgang Menzel. 2017 · 2017
Cited alongside, same era.
Google’s multilingual neural machine translation system: Enabling zero-shot translation
Melvin Johnson, Mike Schuster, Quoc V. Le, Maxim Krikun, Yonghui Wu, Zhifeng Chen, Nikhil Thorat, Fernanda Viégas, Martin Wattenberg, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2017 · 2017
Cited alongside, same era.
MEANT 2.0: Accurate semantic MT evaluation for any output language
Chi-kiu Lo. 2017 · 2017
Cited alongside, same era.
Blend: a novel combined MT metric based on direct assessment — CASICT-DCU submission to WMT17 metrics task
Qingsong Ma, Yvette Graham, Shugen Wang, and Qun Liu. 2017 · 2017
Cited alongside, same era.
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Multilingual whispers: Generating paraphrases with translation
Christian Federmann, Oussama Elachqar, and Chris Quirk. 2019 · 2019
Later among the works it cites.
Findings of the WMT 2019 shared tasks on quality estimation
Erick Fonseca, Lisa Yankovskaya, André F. T. Martins, Mark Fishel, and Christian Federmann. 2019 · 2019
Later among the works it cites.
Meteor++ 2.0: Adopt syntactic level paraphrase knowledge into machine translation evaluation
Yinuo Guo and Junfeng Hu. 2019 · 2019
Later among the works it cites.
Improved lexically constrained decoding for translation and monolingual rewriting
J. Edward Hu, Huda Khayrallah, Ryan Culkin, Patrick Xia, Tongfei Chen, Matt Post, and Benjamin Van Durme. 2019a · 2019
Later among the works it cites.
Gpipe: Efficient training of giant neural networks using pipeline parallelism
Yanping Huang, Youlong Cheng, Ankur Bapna, Orhan Firat, Dehao Chen, Mia Chen, HyoukJoong Lee, Jiquan Ngiam, Quoc V Le, Yonghui Wu, and zhifeng Chen. 2019 · 2019
Later among the works it cites.
Machine translation evaluation using bi-directional entailment
Rakesh Khobragade, Heaven Patel, Anand Namdev, Anish Mishra, and Pushpak Bhattacharyya. 2019 · 2019
Later among the works it cites.
Investigating multilingual NMT representations at scale
Sneha Kudugunta, Ankur Bapna, Isaac Caswell, and Orhan Firat. 2019 · 2019
Later among the works it cites.
YiSi - a unified semantic MT quality evaluation and estimation metric for languages with different levels of available resources
Chi-kiu Lo. 2019 · 2019
Later among the works it cites.
Results of the WMT19 metrics shared task: Segment-level and strong MT systems pose big challenges
Qingsong Ma, Johnny Wei, Ondřej Bojar, and Yvette Graham. 2019 · 2019
Later among the works it cites.
Putting evaluation in context: Contextual embeddings improve machine translation evaluation
Nitika Mathur, Timothy Baldwin, and Trevor Cohn. 2019 · 2019
Later among the works it cites.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Later among the works it cites.
Improving zero-shot translation with language-independent constraints
Ngoc-Quan Pham, Jan Niehues, Thanh-Le Ha, and Alexander Waibel. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
An evaluation of language-agnostic inner-attention-based representations in machine translation
Alessandro Raganato, Raúl Vázquez, Mathias Creutz, and Jörg Tiedemann. 2019 · 2019
Later among the works it cites.
EED: Extended edit distance measure for machine translation
Peter Stanchev, Weiyue Wang, and Hermann Ney. 2019 · 2019
Later among the works it cites.
Measuring semantic abstraction of multilingual NMT with paraphrase recognition and generation tasks
Jörg Tiedemann and Yves Scherrer. 2019 · 2019
Later among the works it cites.
Simple and effective paraphrastic similarity from parallel translations
John Wieting, Kevin Gimpel, Graham Neubig, and Taylor Berg-Kirkpatrick. 2019 · 2019
Later among the works it cites.
Quality estimation and translation metrics via pre-trained word and sentence embeddings
Elizaveta Yankovskaya, Andre Tättar, and Mark Fishel. 2019 · 2019
Later among the works it cites.
Filtering pseudo-references by paraphrasing for automatic evaluation of machine translation
Ryoma Yoshimura, Hiroki Shimanaka, Yukio Matsumura, Hayahide Yamagishi, and Mamoru Komachi. 2019 · 2019
Later among the works it cites.
On the evaluation of machine translation systems trained with back-translation
Sergey Edunov, Myle Ott, Marc’Aurelio Ranzato, and Michael Auli. 2020 · 2020
Closest in time.
Multilingual denoising pre-training for neural machine translation
Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020 · 2020
Closest in time.
Tangled up in BLEU: Reevaluating the evaluation of automatic machine translation evaluation metrics
Nitika Mathur, Timothy Baldwin, and Trevor Cohn. 2020 · 2020
Closest in time.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Closest in time.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Closest in time.