USR: An unsupervised and reference free evaluation metric for dialog generation
Shikib Mehri and Maxine Eskenazi. 2020 · 2020
Later among the works it cites.
Domain robustness in neural machine translation
Mathias Müller, Annette Rios, and Rico Sennrich. 2020 · 2020
Later among the works it cites.
Towards holistic and automatic evaluation of open-domain dialogue generation
Bo Pang, Erik Nijkamp, Wenjuan Han, Linqi Zhou, Yixian Liu, and Kewei Tu. 2020 · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Later among the works it cites.
Recipes for building an open-domain chatbot
Original
Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Kurt Shuster, Eric M. Smith, Y-Lan Boureau, and Jason Weston. 2020 · 2020
Later among the works it cites.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Later among the works it cites.
The dialogue dodecathlon: Open-domain knowledge and image grounded conversational agents
Kurt Shuster, Da Ju, Stephen Roller, Emily Dinan, Y-Lan Boureau, and Jason Weston. 2020 · 2020
Later among the works it cites.
Asking and answering questions to evaluate the factual consistency of summaries
Alex Wang, Kyunghyun Cho, and Mike Lewis. 2020 · 2020
Later among the works it cites.
On exposure bias, hallucination and domain shift in neural machine translation
Chaojun Wang and Rico Sennrich. 2020 · 2020
Later among the works it cites.
Fact-based content weighting for evaluating abstractive summarisation
Xinnuo Xu, Ondřej Dušek, Jingyi Li, Verena Rieser, and Ioannis Konstas. 2020 · 2020
Later among the works it cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Later among the works it cites.
Designing precise and robust dialogue response evaluators
Tianyu Zhao, Divesh Lala, and Tatsuya Kawahara. 2020 · 2020
Later among the works it cites.
Proceedings of the 1st Workshop on Natural Language Generation, Evaluation, and Metrics (GEM 2021) . Association for Computational Linguistics, Online
Antoine Bosselut, Esin Durmus, Varun Prashant Gangal, Sebastian Gehrmann, Yacine Jernite, Laura Perez-Beltrachini, Samira Shaikh, and Wei Xu, editors. 2021 · 2021
Closest in time.
Evaluating groundedness in dialogue systems: The begin benchmark
Original
Nouha Dziri, Hannah Rashkin, Tal Linzen, and David Reitter. 2021 · 2021
Closest in time.
Improving factual consistency of abstractive summarization via question answering
Feng Nan, Cicero Nogueira dos Santos, Henghui Zhu, Patrick Ng, Kathleen McKeown, Ramesh Nallapati, Dejiao Zhang, Zhiguo Wang, Andrew O. Arnold, and Bing Xiang. 2021 · 2021
Closest in time.
Increasing faithfulness in knowledge-grounded dialogue with controllable features
Hannah Rashkin, David Reitter, Gaurav Singh Tomar, and Dipanjan Das. 2021 · 2021
Closest in time.
Questeval: Summarization asks for fact-based evaluation
Original
Thomas Scialom, Paul-Alexis Dray, Gallinari Patrick, Lamprier Sylvain, Piwowarski Benjamin, Staiano Jacopo, and Wang Alex. 2021 · 2021
Closest in time.