A new measure of rank correlation
Maurice G Kendall. 1938 · 1938
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics . 311–318
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries. In Text summarization branches out . 74–81
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Proceedings of the acl workshop on intrinsic and extrinsic evaluation measures for machine translation and/or summarization . 65–72
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Spearman rank correlation
Jerrold H Zar. 2005 · 2005
Earlier work this paper cites.
Phrase-based statistical language generation using graphical models and active learning. In Proceedings of the 48th Annual Meeting of the Association for Computational Linguistics . 1552–1561
François Mairesse, Milica Gasic, Filip Jurcicek, Simon Keizer, Blaise Thomson, Kai Yu, and Steve Young. 2010 · 2010
Earlier work this paper cites.
A guide to appropriate use of correlation coefficient in medical research
Mavuto M Mukaka. 2012 · 2012
Earlier work this paper cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Original
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2018 · 2018
Earlier work this paper cites.
Reliability and learnability of human bandit feedback for sequence-to-sequence reinforcement learning
Original
Julia Kreutzer, Joshua Uyheng, and Stefan Riezler. 2018 · 2018
Earlier work this paper cites.
RankME: Reliable human ratings for natural language generation
Original
Jekaterina Novikova, Ondřej Dušek, and Verena Rieser. 2018 · 2018
Earlier work this paper cites.
Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension
Original
Mike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad, Abdelrahman Mohamed, Omer Levy, Ves Stoyanov, and Luke Zettlemoyer. 2019 · 2019
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Original
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019 · 2019
Earlier work this paper cites.
Evaluation of text generation: A survey
Original
Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020 · 2020
Earlier work this paper cites.
FEQA: A question answering evaluation framework for faithfulness assessment in abstractive summarization
Original
Esin Durmus, He He, and Mona Diab. 2020 · 2020
Earlier work this paper cites.
Unsupervised evaluation of interactive dialog with dialogpt
Original
Shikib Mehri and Maxine Eskenazi. 2020a · 2020
Earlier work this paper cites.