Fetching the paper…
Reading the bibliography…
Automatic evaluation metrics are a crucial component of dialog systems research.
The second conversational intelligence challenge (convai2)
Emily Dinan, Varvara Logacheva, Valentin Malykh, Alexander Miller, Kurt Shuster, Jack Urbanek, Douwe Kiela, Arthur Szlam, Iulian Serban, Ryan Lowe, et al. 2019 · 1902
Earlier work this paper cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019 · 1904
Earlier work this paper cites.
PARADISE: A framework for evaluating spoken dialogue agents
Marilyn A. Walker, Diane J. Litman, Candace A. Kamm, and Alicia Abella. 1997 · 1997
Earlier work this paper cites.
Towards a human-like open-domain chatbot
D. Adiwardana, Minh-Thang Luong, David R. So, J. Hall, Noah Fiedel, R. Thoppilan, Z. Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, and Quoc V. Le. 2020a · 2001
Earlier work this paper cites.
Towards a human-like open-domain chatbot
Daniel Adiwardana, Minh-Thang Luong, David R So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, et al. 2020b · 2001
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Some issues in automatic evaluation of english-hindi mt: more blues for bleu
R Ananthakrishnan, Pushpak Bhattacharyya, M Sasikumar, and Ritesh M Shah. 2007 · 2007
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen. 2010 · 2010
Earlier work this paper cites.
Data-driven response generation in social media
Alan Ritter, Colin Cherry, and William B. Dolan. 2011 · 2011
Earlier work this paper cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean. 2013 · 2013
Earlier work this paper cites.
Learning phrase representations using RNN encoder–decoder for statistical machine translation
Kyunghyun Cho, Bart van Merriënboer, Caglar Gulcehre, Dzmitry Bahdanau, Fethi Bougares, Holger Schwenk, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
deltaBLEU: A discriminative metric for generation tasks with intrinsically diverse targets
Michel Galley, Chris Brockett, Alessandro Sordoni, Yangfeng Ji, Michael Auli, Chris Quirk, Margaret Mitchell, Jianfeng Gao, and Bill Dolan. 2015 · 2015
Earlier work this paper cites.
Ryan Kiros, Yukun Zhu, Ruslan Salakhutdinov, Richard S Zemel, Antonio Torralba, Raquel Urtasun, and Sanja Fidler. 2015 · 2015
Earlier work this paper cites.
The Ubuntu dialogue corpus: A large dataset for research in unstructured multi-turn dialogue systems
Ryan Lowe, Nissan Pow, Iulian Serban, and Joelle Pineau. 2015 · 2015
Earlier work this paper cites.
Yukun Zhu, Ryan Kiros, Rich Zemel, Ruslan Salakhutdinov, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Earlier work this paper cites.
A diversity-promoting objective function for neural conversation models
Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. 2016 · 2016
Earlier work this paper cites.
How NOT to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation
Chia-Wei Liu, Ryan Lowe, Iulian Serban, Mike Noseworthy, Laurent Charlin, and Joelle Pineau. 2016a · 2016
Earlier work this paper cites.
End-to-end conversation modeling track in dstc6
Chiori Hori and Takaaki Hori. 2017 · 2017
Earlier work this paper cites.
DailyDialog: A manually labelled multi-turn dialogue dataset
Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017 · 2017
Earlier work this paper cites.
Towards an automatic Turing test: Learning to evaluate dialogue responses
Ryan Lowe, Michael Noseworthy, Iulian Vlad Serban, Nicolas Angelard-Gontier, Yoshua Bengio, and Joelle Pineau. 2017 · 2017
Cited alongside, same era.
Why we need new evaluation metrics for NLG
Jekaterina Novikova, Ondřej Dušek, Amanda Cercas Curry, and Verena Rieser. 2017 · 2017
Cited alongside, same era.
Topic-based evaluation for conversational bots
Fenfei Guo, A. Metallinou, Chandra Khatri, Anirudh Raju, Anu Venkatesh, and A. Ram. 2018 · 2018
Cited alongside, same era.
A structured review of the validity of BLEU
Ehud Reiter. 2018 · 2018
Cited alongside, same era.
RUSE: Regressor using sentence embeddings for automatic machine translation evaluation
Hiroki Shimanaka, Tomoyuki Kajiwara, and Mamoru Komachi. 2018 · 2018
Cited alongside, same era.
GRADE: Automatic graph-enhanced coherence metric for evaluating open-domain dialogue systems
Lishan Huang, Zheng Ye, Jinghui Qin, Liang Lin, and Xiaodan Liang. 2020 · 2020
Later among the works it cites.
Pone: A novel automatic evaluation metric for open-domain generative dialogue systems
Tian Lan, Xian-Ling Mao, Wei Wei, Xiaoyan Gao, and Heyan Huang. 2020 · 2020
Later among the works it cites.
Cc-news-en: A large english news corpus
Joel Mackenzie, Rodger Benham, Matthias Petri, Johanne R Trippas, J Shane Culpepper, and Alistair Moffat. 2020 · 2020
Later among the works it cites.
Language model transformers as evaluators for open-domain dialogues
Rostislav Nedelchev, Jens Lehmann, and Ricardo Usbeck. 2020 · 2020
Later among the works it cites.
Towards holistic and automatic evaluation of open-domain dialogue generation
Bo Pang, Erik Nijkamp, Wenjuan Han, Linqi Zhou, Yixian Liu, and Kewei Tu. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Elior Sulem, Omri Abend, and Ari Rappoport. 2018 · 2018
Cited alongside, same era.
Ruber: An unsupervised method for automatic evaluation of open-domain dialog systems
Chongyang Tao, Lili Mou, Dongyan Zhao, and Rui Yan. 2018 · 2018
Cited alongside, same era.
On evaluating and comparing conversational agents
Anu Venkatesh, Chandra Khatri, Ashwin Ram, Fenfei Guo, Raefer Gabriel, Ashish Nagar, Rohit Prasad, Ming Cheng, Behnam Hedayatnia, Angeliki Metallinou, et al. 2018 · 2018
Cited alongside, same era.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 · 2018
Cited alongside, same era.
Personalizing dialogue agents: I have a dog, do you have pets too?
Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston. 2018 · 2018
Cited alongside, same era.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Cited alongside, same era.
Better automatic evaluation of open-domain dialogue systems with contextualized embeddings
Sarik Ghazarian, Johnny Wei, Aram Galstyan, and Nanyun Peng. 2019 · 2019
Cited alongside, same era.
Deconstruct to reconstruct a configurable evaluation metric for open-domain dialogue systems
Vitou Phy, Yang Zhao, and Akiko Aizawa. 2020 · 2020
Later among the works it cites.
Exploring the limits of transfer learning with a unified text-to-text transformer
Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020 · 2020
Later among the works it cites.
Improving dialog evaluation with a multi-reference adversarial dataset and large scale pretraining
Ananya B. Sai, Akash Kumar Mohankumar, Siddhartha Arora, and Mitesh M. Khapra. 2020 · 2020
Later among the works it cites.
BLEURT: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Later among the works it cites.
Learning an unreferenced metric for online dialogue evaluation
Koustuv Sinha, Prasanna Parthasarathi, Jasmine Wang, Ryan Lowe, William L. Hamilton, and Joelle Pineau. 2020 · 2020
Later among the works it cites.
uBLEU: Uncertainty-aware automatic evaluation method for open-domain dialogue systems
Tsuta Yuma, Naoki Yoshinaga, and Masashi Toyoda. 2020 · 2020
Later among the works it cites.
Dialogpt: Large-scale generative pre-training for conversational response generation
Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan. 2020 · 2020
Later among the works it cites.
Learning to compare for better training and evaluation of open domain natural language generation models
Wangchunshu Zhou and Ke Xu. 2020 · 2020
Later among the works it cites.
Survey on evaluation methods for dialogue systems
Jan Deriu, Alvaro Rodrigo, Arantxa Otegi, Guillermo Echegoyen, Sophie Rosset, Eneko Agirre, and Mark Cieliebak. 2021 · 2021
Closest in time.
Overview of the ninth dialog system technology challenge: Dstc9
Chulaka Gunasekara, Seokhwan Kim, Luis Fernando D’Haro, Abhinav Rastogi, Yun-Nung Chen, Mihail Eric, Behnam Hedayatnia, Karthik Gopalakrishnan, Yang Liu, Chao-Wei Huang, et al. 2021 · 2021
Closest in time.
Data-questeval: A referenceless metric for data to text semantic evaluation
Cl’ement Rebuffel, Thomas Scialom, Laure Soulier, Benjamin Piwowarski, Sylvain Lamprier, Jacopo Staiano, Geoffrey Scoutheeten, and P. Gallinari. 2021 · 2021
Closest in time.
Automatic evaluation of non-task oriented dialog systems by using sentence embeddings projections and their dynamics
Mario Rodríguez-Cantelar, Luis Fernando D’Haro, and Fernando Matía. 2021 · 2021
Closest in time.
Questeval: Summarization asks for fact-based evaluation
Thomas Scialom, Paul-Alexis Dray, Patrick Gallinari, Sylvain Lamprier, Benjamin Piwowarski, Jacopo Staiano, and Alex Wang. 2021 · 2021
Closest in time.
Assessing dialogue systems with distribution distances
Jiannan Xiang, Yahui Liu, Deng Cai, Huayang Li, Defu Lian, and Lemao Liu. 2021 · 2021
Closest in time.
Dynaeval: Unifying turn and dialogue level evaluation
Chen Zhang, Yiming Chen, Luis Fernando D’Haro, Yan Zhang, Thomas Friedrichs, Grandee Lee, and Haizhou Li. 2021a · 2021
Closest in time.