Fetching the paper…
Reading the bibliography…
As machine translation (MT) metrics improve their correlation with human judgement every year, it is crucial to understand the limitations of such metrics at the segment level.
Using test suites in evaluation of machine translation systems
Margaret King and Kirsten Falkedal. 1990 · 1990
Earlier work this paper cites.
WordNet: A lexical database for English
George A. Miller. 1994 · 1994
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Europarl: A parallel corpus for statistical machine translation
Philipp Koehn. 2005 · 2005
Earlier work this paper cites.
Unbounded dependency recovery for parser evaluation
Laura Rimell, Stephen Clark, and Mark Steedman. 2009 · 2009
Earlier work this paper cites.
Babelnet: The automatic construction, evaluation and application of a wide-coverage multilingual semantic network
Roberto Navigli and Simone Paolo Ponzetto. 2012 · 2012
Earlier work this paper cites.
Adversarial evaluation for models of natural language
Noah A. Smith. 2012 · 2012
Earlier work this paper cites.
Multidimensional quality metrics (mqm): A framework for declaring and describing translation quality metrics
Arle Lommel, Aljoscha Burchardt, and Hans Uszkoreit. 2014 · 2014
Earlier work this paper cites.
Neural versus phrase-based machine translation quality: a case study
Luisa Bentivogli, Arianna Bisazza, Mauro Cettolo, and Marcello Federico. 2016 · 2016
Earlier work this paper cites.
PROTEST: A test suite for evaluating pronouns in machine translation
Liane Guillou and Christian Hardmeier. 2016 · 2016
Earlier work this paper cites.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Is neural machine translation the new state of the art?
Sheila Castilho, Joss Moorkens, Federico Gaspari, Iacer Calixto, John Tinsley, and Andy Way. 2017 · 2017
Earlier work this paper cites.
A challenge set approach to evaluating machine translation
Pierre Isabelle, Colin Cherry, and George Foster. 2017 · 2017
Earlier work this paper cites.
Improving discourse relation projection to build discourse annotated corpora
Majid Laali and Leila Kosseim. 2017 · 2017
Earlier work this paper cites.
BIBI system description: Building with CNNs and breaking with deep reinforcement learning
Yitong Li, Trevor Cohn, and Timothy Baldwin. 2017 · 2017
Earlier work this paper cites.
Breaking NLP: Using morphosyntax, semantics, pragmatics and world knowledge to fool sentiment analysis systems
Taylor Mahler, Willy Cheung, Micha Elsner, David King, Marie-Catherine de Marneffe, Cory Shain, Symon Stevens-Guille, and Michael White. 2017 · 2017
Earlier work this paper cites.
chrF++: words helping character n-grams
Maja Popović. 2017 · 2017
Earlier work this paper cites.
Social bias in elicited natural language inferences
Rachel Rudinger, Chandler May, and Benjamin Van Durme. 2017 · 2017
Earlier work this paper cites.
Breaking sentiment analysis of movie reviews
Ieva Staliūnaitė and Ben Bonfil. 2017 · 2017
Earlier work this paper cites.
A multifaceted evaluation of neural versus phrase-based machine translation for 9 language directions
Antonio Toral and Víctor M. Sánchez-Cartagena. 2017 · 2017
Earlier work this paper cites.
Attention is All you Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Earlier work this paper cites.
Fine-grained evaluation of quality estimation for machine translation based on a linguistically motivated test suite
Eleftherios Avramidis, Vivien Macketanz, Arle Lommel, and Hans Uszkoreit. 2018 · 2018
Earlier work this paper cites.
XNLI: Evaluating cross-lingual sentence representations
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018 · 2018
Earlier work this paper cites.
A pronoun test suite evaluation of the English–German MT systems at WMT 2018
Liane Guillou, Christian Hardmeier, Ekaterina Lapshinova-Koltunski, and Sharid Loáiciga. 2018 · 2018
Cited alongside, same era.
Looking beyond the surface: A challenge set for reading comprehension over multiple sentences
Daniel Khashabi, Snigdha Chaturvedi, Michael Roth, Shyam Upadhyay, and Dan Roth. 2018 · 2018
Cited alongside, same era.
SentencePiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Cited alongside, same era.
ParCorFull: a parallel corpus annotated with full coreference
Ekaterina Lapshinova-Koltunski, Christian Hardmeier, and Pauline Krielke. 2018 · 2018
Cited alongside, same era.
Gender bias in coreference resolution: Evaluation and debiasing methods
Jieyu Zhao, Tianlu Wang, Mark Yatskar, Vicente Ordonez, and Kai-Wei Chang. 2018 · 2018
Cited alongside, same era.
Bertscore: Evaluating text generation with BERT
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Later among the works it cites.
Wino-X: Multilingual Winograd schemas for commonsense reasoning and coreference resolution
Denis Emelin and Rico Sennrich. 2021 · 2021
Later among the works it cites.
Beyond english-centric multilingual machine translation
Angela Fan, Shruti Bhosale, Holger Schwenk, Zhiyi Ma, Ahmed El-Kishky, Siddharth Goyal, Mandeep Baines, Onur Celebi, Guillaume Wenzek, Vishrav Chaudhary, Naman Goyal, Tom Birch, Vitaliy Liptchinsky, Sergey Edunov, Michael Auli, and Armand Joulin. 2021 · 2021
Later among the works it cites.
A fine-grained analysis of BERTScore
Michael Hanna and Ondřej Bojar. 2021 · 2021
Later among the works it cites.
To ship or not to ship: An extensive evaluation of automatic metrics for machine translation
Tom Kocmi, Christian Federmann, Roman Grundkiewicz, Marcin Junczys-Dowmunt, Hitokazu Matsushita, and Arul Menezes. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
YiSi - a unified semantic MT quality evaluation and estimation metric for languages with different levels of available resources
Chi-kiu Lo. 2019 · 2019
Cited alongside, same era.
Non-entailed subsequences as a challenge for natural language inference
Richard T McCoy and Tal Linzen. 2019 · 2019
Cited alongside, same era.
Challenge test sets for MT evaluation
Maja Popović and Sheila Castilho. 2019 · 2019
Cited alongside, same era.
The MuCoW test suite at WMT 2019: Automatically harvested multilingual contrastive word sense disambiguation test sets for machine translation
Alessandro Raganato, Yves Scherrer, and Jörg Tiedemann. 2019 · 2019
Cited alongside, same era.
Sentence-BERT: Sentence embeddings using Siamese BERT-networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Cited alongside, same era.
Evaluating gender bias in machine translation
Gabriel Stanovsky, Noah A. Smith, and Luke Zettlemoyer. 2019 · 2019
Cited alongside, same era.
NoiseQA: Challenge set evaluation for user-centric question answering
Abhilasha Ravichander, Siddharth Dalmia, Maria Ryskina, Florian Metze, Eduard Hovy, and Alan W Black. 2021 · 2021
Later among the works it cites.
Fancy: A diagnostic data-set for nli models
Guido Rocchietti, Flavia Achena, Giuseppe Marziano, Sara Salaris, and Alessandro Lenci. 2021 · 2021
Later among the works it cites.
Masked language modeling and the distributional hypothesis: Order word matters pre-training for little
Koustuv Sinha, Robin Jia, Dieuwke Hupkes, Joelle Pineau, Adina Williams, and Douwe Kiela. 2021 · 2021
Later among the works it cites.
Contrastive conditioning for assessing disambiguation in MT: A case study of distilled bias
Jannis Vamvas and Rico Sennrich. 2021 · 2021
Later among the works it cites.
Understanding the societal impacts of machine translation: a critical review of the literature on medical and legal use cases
Lucas Nunes Vieira, Minako O’Hagan, and Carol O’Sullivan. 2021 · 2021
Later among the works it cites.
PIE: A parallel idiomatic expression corpus for idiomatic sentence generation and paraphrasing
Jianing Zhou, Hongyu Gong, and Suma Bhat. 2021 · 2021
Later among the works it cites.
Chantal Amrhein and Rico Sennrich. 2022 · 2022
Closest in time.
Can transformer be too compositional? analysing idiom processing in neural machine translation
Verna Dankers, Christopher Lucas, and Ivan Titov. 2022 · 2022
Closest in time.
The Flores-101 evaluation benchmark for low-resource and multilingual machine translation
Naman Goyal, Cynthia Gao, Vishrav Chaudhary, Peng-Jen Chen, Guillaume Wenzek, Da Ju, Sanjana Krishnan, Marc’Aurelio Ranzato, Francisco Guzmán, and Angela Fan. 2022 · 2022
Closest in time.
MS-COMET: More and Better Human Judgements Improve Metric Performance
Tom Kocmi, Hitokazu Matsushita, and Christian Federmann. 2022 · 2022
Closest in time.
Partial Could Be Better Than Whole: HW-TSC 2022 Submission for the Metrics Shared Task
Yilun Liu, Xiaosong Qiao, Zhanglin Wu, Su Chang, Min Zhang, Yanqing Zhao, shimin tao Song Peng, Hao Yang, Ying Qin, Jiaxin Guo, Minghan Wang, Yinglu Li, Peng Li, and Xiaofeng Zhao. 2022 · 2022
Closest in time.
No language left behind: Scaling human-centered machine translation
NLLB Team, Marta R. Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, Anna Sun, Skyler Wang, Guillaume Wenzek, Al Youngblood, Bapi Akula, Loic Barrault, Gabriel Mejia Gonzalez, Prangthip Hansanti, John Hoffman, Semarley Jarrett, Kaushik Ram Sadagopan, Dirk Rowe, Shannon Spruit, Chau Tran, Pierre Andrews, Necip Fazil Ayan, Shruti Bhosale, Sergey Edunov, Angela Fan, Cynthia Gao, Vedanuj Goswami, Francisco Guzmán, Philipp Koehn, Alexandre Mourachko, Christophe Ropers, Safiyyah Saleem, Holger Schwenk, and Jeff Wang. 2022 · 2022
Closest in time.
Machine Translation Evaluation as a Sequence Tagging Problem
Stefano Perrella, Lorenzo Proietti, Alessandro Scirè, Niccolò Campolungo, and Roberto Navigli. 2022 · 2022
Closest in time.
COMET-22: Unbabel-IST 2022 Submission for the Metrics Shared Task
Ricardo Rei, José G. C. de Souza, Duarte Alves, Chrysoula Zerva, Ana C Farinha, Taisiya Glushkova, Alon Lavie, Luisa Coheur, and André F. T. Martins. 2022 · 2022
Closest in time.
As little as possible, as much as necessary: Detecting over- and undertranslations with contrastive conditioning
Jannis Vamvas and Rico Sennrich. 2022 · 2022
Closest in time.
Alibaba-Translate China’s Submission for WMT2022 Metrics Shared Task
Yu Wan, Keqin Bao, Dayiheng Liu, Baosong Yang, Derek F. Wong, Lidia S. Chao, Wenqiang Lei, and Jun Xie. 2022 · 2022
Closest in time.
Adversarial examples for evaluating reading comprehension systems
Robin Jia and Percy Liang. 2017 · 2031
Closest in time.