Fetching the paper…
Reading the bibliography…
A major challenge in the field of Text Generation is evaluation because we lack a sound theory that can be leveraged to extract guidelines for evaluation campaigns.
Bleu: a Method for Automatic Evaluation of Machine Translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Re-evaluating the role of Bleu in machine translation research
Chris Callison-Burch, Miles Osborne, and Philipp Koehn. 2006 · 2006
Earlier work this paper cites.
Evaluation of text generation: A survey
Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020 · 2006
Earlier work this paper cites.
Multidimensional quality metrics (mqm): A framework for declaring and describing translation quality metrics
Arle Lommel, Hans Uszkoreit, and Aljoscha Burchardt. 2014 · 2014
Earlier work this paper cites.
Why we need new evaluation metrics for NLG
Jekaterina Novikova, Ondřej Dušek, Amanda Cercas Curry, and Verena Rieser. 2017 · 2017
Earlier work this paper cites.
The price of debiasing automatic metrics in natural language evalaution
Arun Chaganty, Stephen Mussmann, and Percy Liang. 2018 · 2018
Earlier work this paper cites.
Rankme: Reliable human ratings for natural language generation
Jekaterina Novikova, Ondřej Dušek, and Verena Rieser. 2018 · 2018
Earlier work this paper cites.
Comparing bayesian models of annotation
Silviu Paun, Bob Carpenter, Jon Chamberlain, Dirk Hovy, Udo Kruschwitz, and Massimo Poesio. 2018 · 2018
Earlier work this paper cites.
A call for clarity in reporting BLEU scores
Matt Post. 2018 · 2018
Earlier work this paper cites.
The use of rating and Likert scales in natural language generation human evaluation tasks: A review and some recommendations
Jacopo Amidei, Paul Piwek, and Alistair Willis. 2019 · 2019
Earlier work this paper cites.
Unifying human and statistical evaluation for natural language generation
Tatsunori B. Hashimoto, Hugh Zhang, and Percy Liang. 2019 · 2019
Cited alongside, same era.
Towards best experiment design for evaluating dialogue system output
Sashank Santhanam and Samira Shaikh. 2019 · 2019
Cited alongside, same era.
Best practices for the human evaluation of automatically generated text
Chris Van Der Lee, Albert Gatt, Emiel Van Miltenburg, Sander Wubben, and Emiel Krahmer. 2019 · 2019
Cited alongside, same era.
With little power comes great responsibility
Dallas Card, Peter Henderson, Urvashi Khandelwal, Robin Jia, Kyle Mahowald, and Dan Jurafsky. 2020 · 2020
Cited alongside, same era.
Spot the bot: A robust and efficient framework for the evaluation of conversational dialogue systems
Jan Deriu, Don Tuggener, Pius von Däniken, Jon Ander Campos, Alvaro Rodrigo, Thiziri Belkacem, Aitor Soroa, Eneko Agirre, and Mark Cieliebak. 2020 · 2020
Cited alongside, same era.
The University of Edinburgh’s English-German and English-Hausa submissions to the WMT21 news translation task
Pinzhen Chen, Jindřich Helcl, Ulrich Germann, Laurie Burchell, Nikolay Bogoychev, Antonio Valerio Miceli Barone, Jonas Waldendorf, Alexandra Birch, and Kenneth Heafield. 2021 · 2021
Later among the works it cites.
All that’s ‘human’ is not gold: Evaluating human evaluation of generated text
Elizabeth Clark, Tal August, Sofia Serrano, Nikita Haduong, Suchin Gururangan, and Noah A. Smith. 2021 · 2021
Later among the works it cites.
Survey on evaluation methods for dialogue systems
Jan Deriu, Alvaro Rodrigo, Arantxa Otegi, Guillermo Echegoyen, Sophie Rosset, Eneko Agirre, and Mark Cieliebak. 2021 · 2021
Later among the works it cites.
The volctrans GLAT system: Non-autoregressive translation meets WMT21
Lihua Qian, Yi Zhou, Zaixiang Zheng, Yaoming Zhu, Zehui Lin, Jiangtao Feng, Shanbo Cheng, Lei Li, Mingxuan Wang, and Hao Zhou. 2021 · 2021
Later among the works it cites.
Recipes for building an open-domain chatbot
Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Eric Michael Smith, Y-Lan Boureau, and Jason Weston. 2021 · 2021
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
The second conversational intelligence challenge (convai2)
Emily Dinan, Varvara Logacheva, Valentin Malykh, Alexander Miller, Kurt Shuster, Jack Urbanek, Douwe Kiela, Arthur Szlam, Iulian Serban, Ryan Lowe, et al. 2020 · 2020
Cited alongside, same era.
USR: An unsupervised and reference free evaluation metric for dialog generation
Shikib Mehri and Maxine Eskenazi. 2020 · 2020
Cited alongside, same era.
COMET: A neural framework for MT evaluation
Ricardo Rei, Craig Stewart, Ana C Farinha, and Alon Lavie. 2020 · 2020
Cited alongside, same era.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Cited alongside, same era.
The reprogen shared task on reproducibility of human evaluations in nlg: Overview and results
Anja Belz, Anastasia Shimorina, Shubham Agarwal, and Ehud Reiter. 2021 · 2021
Cited alongside, same era.
Experts, errors, and context: A large-scale study of human evaluation for machine translation
Markus Freitag, George Foster, David Grangier, Viresh Ratnakar, Qijun Tan, and Wolfgang Macherey. 2021a
Cited in the paper.
Results of the WMT21 metrics shared task: Evaluating metrics with expert-based human evaluations on TED and news domain
Markus Freitag, Ricardo Rei, Nitika Mathur, Chi-kiu Lo, Craig Stewart, George Foster, Alon Lavie, and Ondřej Bojar. 2021b
Cited in the paper.
Later among the works it cites.
Facebook AI’s WMT21 news translation task submission
Chau Tran, Shruti Bhosale, James Cross, Philipp Koehn, Sergey Edunov, and Angela Fan. 2021 · 2021
Later among the works it cites.
HW-TSC’s participation in the WMT 2021 news translation shared task
Daimeng Wei, Zongyao Li, Zhanglin Wu, Zhengzhe Yu, Xiaoyu Chen, Hengchao Shang, Jiaxin Guo, Minghan Wang, Lizhi Lei, Min Zhang, Hao Yang, and Ying Qin. 2021 · 2021
Later among the works it cites.
How human is human evaluation? Improving the gold standard for NLG with utility theory
Kawin Ethayarajh and Dan Jurafsky. 2022 · 2022
Closest in time.
Active evaluation: Efficient nlg evaluation with few pairwise comparisons
Akash Kumar Mohankumar and Mitesh M Khapra. 2022 · 2022
Closest in time.
A survey of evaluation metrics used for nlg systems
Ananya B Sai, Akash Kumar Mohankumar, and Mitesh M Khapra. 2022 · 2022
Closest in time.