Fetching the paper…
Reading the bibliography…
The advent and fast development of neural networks have revolutionized the research on dialogue systems and subsequently have triggered various challenges regarding their automatic evaluation.
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 1901
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Chameleons in imagined conversations: A new approach to understanding coordination of linguistic style in dialogs
Cristian Danescu-Niculescu-Mizil and Lillian Lee. 2011 · 2011
Earlier work this paper cites.
Movie-DiC: a movie dialogue corpus for research and development
Rafael E. Banchs. 2012 · 2012
Earlier work this paper cites.
Neural responding machine for short-text conversation
Lifeng Shang, Zhengdong Lu, and Hang Li. 2015 · 2015
Earlier work this paper cites.
Intestinal microbiota is influenced by gender and body mass index
Carmen Haro, Oriol A Rangel-Zúñiga, Juan F Alcalá-Díaz, Francisco Gómez-Delgado, Pablo Pérez-Martínez, Javier Delgado-Lista, Gracia M Quintana-Navarro, Blanca B Landa, Juan A Navas-Cortés, Manuel Tena-Sempere, et al. 2016 · 2016
Earlier work this paper cites.
The dialogue breakdown detection challenge: Task description, datasets, and evaluation metrics
Ryuichiro Higashinaka, Kotaro Funakoshi, Yuka Kobayashi, and Michimasa Inaba. 2016 · 2016
Earlier work this paper cites.
How NOT to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation
Chia-Wei Liu, Ryan Lowe, Iulian Serban, Mike Noseworthy, Laurent Charlin, and Joelle Pineau. 2016 · 2016
Earlier work this paper cites.
DailyDialog: A manually labelled multi-turn dialogue dataset
Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017 · 2017
Earlier work this paper cites.
Sequential matching network: A new architecture for multi-turn response selection in retrieval-based chatbots
Yu Wu, Wei Wu, Chen Xing, Ming Zhou, and Zhoujun Li. 2017 · 2017
Earlier work this paper cites.
EmotionLines: An emotion corpus of multi-party conversations
Chao-Chun Hsu, Sheng-Yeh Chen, Chuan-Chun Kuo, Ting-Hao Huang, and Lun-Wei Ku. 2018 · 2018
Earlier work this paper cites.
Towards exploiting background knowledge for building conversation systems
Nikita Moghe, Siddhartha Arora, Suman Banerjee, and Mitesh M. Khapra. 2018 · 2018
Earlier work this paper cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, Ilya Sutskever, et al. 2018 · 2018
Earlier work this paper cites.
Emotional chatting machine: Emotional conversation generation with internal and external memory
Hao Zhou, Minlie Huang, Tianyang Zhang, Xiaoyan Zhu, and Bing Liu. 2018a · 2018
Earlier work this paper cites.
A dataset for document grounded conversations
Kangyan Zhou, Shrimai Prabhumoye, and Alan W Black. 2018c · 2018
Earlier work this paper cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Earlier work this paper cites.
Wizard of wikipedia: Knowledge-powered conversational agents
Emily Dinan, Stephen Roller, Kurt Shuster, Angela Fan, Michael Auli, and Jason Weston. 2019 · 2019
Earlier work this paper cites.
Grounded response generation task at dstc7
Michel Galley, Chris Brockett, Xiang Gao, Jianfeng Gao, and Bill Dolan. 2019 · 2019
Cited alongside, same era.
Topical-Chat: Towards Knowledge-Grounded Open-Domain Conversations
Karthik Gopalakrishnan, Behnam Hedayatnia, Qinlang Chen, Anna Gottardi, Sanjeev Kwatra, Anu Venkatesh, Raefer Gabriel, and Dilek Hakkani-Tür. 2019 · 2019
Cited alongside, same era.
Investigating evaluation of open-domain dialogue systems with human generated multiple references
Prakhar Gupta, Shikib Mehri, Tiancheng Zhao, Amy Pavel, Maxine Eskenazi, and Jeffrey P. Bigham. 2019 · 2019
Cited alongside, same era.
Multi-domain task-completion dialog challenge
S Lee, H Schulz, A Atkinson, J Gao, K Suleman, L El Asri, M Adada, M Huang, S Sharma, W Tay, et al. 2019 · 2019
Cited alongside, same era.
MELD: A multimodal multi-party dataset for emotion recognition in conversations
Soujanya Poria, Devamanyu Hazarika, Navonil Majumder, Gautam Naik, Erik Cambria, and Rada Mihalcea. 2019 · 2019
Cited alongside, same era.
Sentimental liar: Extended corpus and deep learning models for fake claim classification
Bibek Upadhayay and Vahid Behzadan. 2020 · 2020
Later among the works it cites.
A large-scale chinese short-text conversation dataset
Yida Wang, Pei Ke, Yinhe Zheng, Kaili Huang, Yong Jiang, Xiaoyan Zhu, and Minlie Huang. 2020a · 2020
Later among the works it cites.
DIALOGPT : Large-scale generative pre-training for conversational response generation
Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan. 2020 · 2020
Later among the works it cites.
Designing precise and robust dialogue response evaluators
Tianyu Zhao, Divesh Lala, and Tatsuya Kawahara. 2020 · 2020
Later among the works it cites.
Parrot: Paraphrase generation for nlu
Prithiviraj Damodaran. 2021 · 2021
Later among the works it cites.
Survey on evaluation methods for dialogue systems
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019 · 2019
Cited alongside, same era.
Towards empathetic open-domain conversation models: A new benchmark and dataset
Hannah Rashkin, Eric Michael Smith, Margaret Li, and Y-Lan Boureau. 2019 · 2019
Cited alongside, same era.
ChatEval: A tool for chatbot evaluation
João Sedoc, Daphne Ippolito, Arun Kirubarajan, Jai Thirani, Lyle Ungar, and Chris Callison-Burch. 2019 · 2019
Cited alongside, same era.
What makes a good conversation? how controllable attributes affect human judgments
Abigail See, Stephen Roller, Douwe Kiela, and Jason Weston. 2019 · 2019
Cited alongside, same era.
Towards a human-like open-domain chatbot
Daniel Adiwardana, Minh-Thang Luong, David R. So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, and Quoc V. Le. 2020 · 2020
Cited alongside, same era.
Is this dialogue coherent? learning from dialogue acts and entities
Alessandra Cervone and Giuseppe Riccardi. 2020 · 2020
Cited alongside, same era.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2020 · 2020
Cited alongside, same era.
Jan Deriu, Alvaro Rodrigo, Arantxa Otegi, Guillermo Echegoyen, Sophie Rosset, Eneko Agirre, and Mark Cieliebak. 2021 · 2021
Later among the works it cites.
A comprehensive assessment of dialog evaluation metrics
Yi-Ting Yeh, Maxine Eskenazi, and Shikib Mehri. 2021 · 2021
Later among the works it cites.
Deep am-fm: Toolkit for automatic dialogue evaluation
Chen Zhang, Luis Fernando D’Haro, Rafael E Banchs, Thomas Friedrichs, and Haizhou Li. 2021 · 2021
Later among the works it cites.
PLATO-XL: Exploring the large-scale pre-training of dialogue generation
Siqi Bao, Huang He, Fan Wang, Hua Wu, Haifeng Wang, Wenquan Wu, Zhihua Wu, Zhen Guo, Hua Lu, Xinxian Huang, Xin Tian, Xinchao Xu, Yingzhan Lin, and Zheng-Yu Niu. 2022 · 2022
Later among the works it cites.
Shikib Mehri, Jinho Choi, Luis Fernando D’Haro, Jan Deriu, Maxine Eskenazi, Milica Gasic, Kallirroi Georgila, Dilek Hakkani-Tur, Zekang Li, Verena Rieser, Samira Shaikh, David Traum, Yi-Ting Yeh, Zhou Yu, Yizhe Zhang, and Chen Zhang. 2022 · 2022
Later among the works it cites.
Blenderbot 3: a deployed conversational agent that continually learns to responsibly engage
Kurt Shuster, Jing Xu, Mojtaba Komeili, Da Ju, Eric Michael Smith, Stephen Roller, Megan Ung, Moya Chen, Kushal Arora, Joshua Lane, Morteza Behrooz, William Ngan, Spencer Poff, Naman Goyal, Arthur Szlam, Y-Lan Boureau, Melanie Kambadur, and Jason Weston. 2022 · 2022
Later among the works it cites.
BERTScore is unfair: On social bias in language model-based metrics for text generation
Tianxiang Sun, Junliang He, Xipeng Qiu, and Xuanjing Huang. 2022 · 2022
Later among the works it cites.
EnDex: Evaluation of dialogue engagingness at scale
Guangxuan Xu, Ruibo Liu, Fabrice Harel-Canada, Nischal Reddy Chandra, and Nanyun Peng. 2022 · 2022
Later among the works it cites.
MDD-Eval: Self-training on augmented data for multi-domain dialogue evaluation
Chen Zhang, Luis Fernando D’Haro, Thomas Friedrichs, and Haizhou Li. 2022a · 2022
Later among the works it cites.
FineD-eval: Fine-grained automatic dialogue-level evaluation
Chen Zhang, Luis Fernando D’Haro, Qiquan Zhang, Thomas Friedrichs, and Haizhou Li. 2022b · 2022
Later among the works it cites.
Automatic evaluation and moderation of open-domain dialogue systems
Chen Zhang, João Sedoc, Luis Fernando D’Haro, Rafael Banchs, and Alexander Rudnicky. 2022c · 2022
Later among the works it cites.
Human-centered metrics for dialog system evaluation
Salvatore Giorgi, Shreya Havaldar, Farhan Ahmed, Zuhaib Akhtar, Shalaka Vaidya, Gary Pan, Lyle H. Ungar, H. Andrew Schwartz, and Joao Sedoc. 2023 · 2023
Closest in time.
OpenAI. 2023 · 2023
Closest in time.