Fetching the paper…
Reading the bibliography…
Despite tremendous advancements in dialogue systems, stable evaluation still requires human judgments producing notoriously high-variance metrics due to their inherent subjectivity.
Towards a Human-like Open-Domain Chatbot
Daniel Adiwardana, Minh-Thang Luong, David R. So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, and Quoc V. Le. 2020 · 2001
Earlier work this paper cites.
The effects of affective interventions in human–computer interaction
Timo Partala and Veikko Surakka. 2004 · 2004
Earlier work this paper cites.
The Empathic Companion: A Character-Based Interface That Addresses Users’ Affective States
Helmut Prendinger and Mitsuru Ishizuka. 2005 · 2005
Earlier work this paper cites.
Statistical theories of mental test scores
Frederic M Lord and Melvin R Novick. 2008 · 2008
Earlier work this paper cites.
An Evaluation Protocol for Generative Conversational Systems
Seolhwa Lee, Heuiseok Lim, and João Sedoc. 2020 · 2010
Earlier work this paper cites.
How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation
Chia-Wei Liu, Ryan Lowe, Iulian Serban, Mike Noseworthy, Laurent Charlin, and Joelle Pineau. 2016 · 2016
Earlier work this paper cites.
ParlAI: A Dialog Research Software Platform
Alexander Miller, Will Feng, Dhruv Batra, Antoine Bordes, Adam Fisch, Jiasen Lu, Devi Parikh, and Jason Weston. 2017 · 2017
Earlier work this paper cites.
Advancing the State of the Art in Open Domain Dialog Systems through the Alexa Prize
Chandra Khatri, Behnam Hedayatnia, Anu Venkatesh, Jeff Nunn, Yi Pan, Qing Liu, Han Song, Anna Gottardi, Sanjeev Kwatra, Sanju Pancholi, Ming Cheng, Qinglang Chen, Lauren Stubel, Karthik Gopalakrishnan, Kate Bland, Raefer Gabriel, Arindam Mandal, Dilek Hakkani-Tur, Gene Hwang, Nate Michel, Eric King, and Rohit Prasad. 2018 · 2018
Earlier work this paper cites.
Conversational AI: The Science Behind the Alexa Prize
Ashwin Ram, Rohit Prasad, Chandra Khatri, Anu Venkatesh, Raefer Gabriel, Qing Liu, Jeff Nunn, Behnam Hedayatnia, Ming Cheng, Ashish Nagar, Eric King, Kate Bland, Amanda Wartick, Yi Pan, Han Song, Sk Jayadevan, Gene Hwang, and Art Pettigrue. 2018 · 2018
Earlier work this paper cites.
Wizard of Wikipedia: Knowledge-Powered Conversational Agents
Emily Dinan, Stephen Roller, Kurt Shuster, Angela Fan, Michael Auli, and Jason Weston. 2019 · 2019
Earlier work this paper cites.
Topical-Chat: Towards Knowledge-Grounded Open-Domain Conversations
Karthik Gopalakrishnan, Behnam Hedayatnia, Qinlang Chen, Anna Gottardi, Sanjeev Kwatra, Anu Venkatesh, Raefer Gabriel, and Dilek Hakkani-Tür. 2019 · 2019
Earlier work this paper cites.
Towards Empathetic Open-domain Conversation Models: A New Benchmark and Dataset
Hannah Rashkin, Eric Michael Smith, Margaret Li, and Y-Lan Boureau. 2019 · 2019
Earlier work this paper cites.
Re-Evaluating ADEM: A Deeper Look at Scoring Dialogue Responses
Ananya B. Sai, Mithun Das Gupta, Mitesh M. Khapra, and Mukundhan Srinivasan. 2019 · 2019
Earlier work this paper cites.
Generating Responses with a Specific Emotion in Dialog
Zhenqiao Song, Xiaoqing Zheng, Lu Liu, Mu Xu, and Xuanjing Huang. 2019 · 2019
Earlier work this paper cites.
Dialogue Natural Language Inference
Sean Welleck, Jason Weston, Arthur Szlam, and Kyunghyun Cho. 2019 · 2019
Earlier work this paper cites.
Bridging the Gap between Prior and Posterior Knowledge Selection for Knowledge-Grounded Dialogue Generation
Xiuyi Chen, Fandong Meng, Peng Li, Feilong Chen, Shuang Xu, Bo Xu, and Jie Zhou. 2020 · 2020
Earlier work this paper cites.
Spot The Bot: A Robust and Efficient Framework for the Evaluation of Conversational Dialogue Systems
Jan Deriu, Don Tuggener, Pius von Däniken, Jon Ander Campos, Alvaro Rodrigo, Thiziri Belkacem, Aitor Soroa, Eneko Agirre, and Mark Cieliebak. 2020 · 2020
Earlier work this paper cites.
Towards Unified Dialogue System Evaluation: A Comprehensive Analysis of Current Evaluation Protocols
Sarah E. Finch and Jinho D. Choi. 2020 · 2020
Earlier work this paper cites.
Emora: An Inquisitive Social Chatbot Who Cares For You
Sarah E. Finch, James D. Finch, Ali Ahmadvand, Ingyu, Choi, Xiangjue Dong, Ruixiang Qi, Harshita Sahijwani, Sergey Volokhin, Zihan Wang, Zihao Wang, and Jinho D. Choi. 2020 · 2020
Earlier work this paper cites.
Dense passage retrieval for open-domain question answering
Vladimir Karpukhin, Barlas Oguz, Sewon Min, Patrick Lewis, Ledell Wu, Sergey Edunov, Danqi Chen, and Wen-tau Yih. 2020 · 2020
Earlier work this paper cites.
Will I Sound Like Me? Improving Persona Consistency in Dialogues through Pragmatic Self-Consciousness
Hyunwoo Kim, Byeongchang Kim, and Gunhee Kim. 2020 · 2020
Cited alongside, same era.
Don’t Say That! Making Inconsistent Dialogue Unlikely with Unlikelihood Training
Margaret Li, Stephen Roller, Ilia Kulikov, Sean Welleck, Y-Lan Boureau, Kyunghyun Cho, and Jason Weston. 2020 · 2020
Cited alongside, same era.
MIME: MIMicking Emotions for Empathetic Response Generation
Navonil Majumder, Pengfei Hong, Shanshan Peng, Jiankun Lu, Deepanway Ghosal, Alexander Gelbukh, Rada Mihalcea, and Soujanya Poria. 2020 · 2020
Cited alongside, same era.
DukeNet: A Dual Knowledge Interaction Network for Knowledge-Grounded Conversation
Chuan Meng, Pengjie Ren, Zhumin Chen, Weiwei Sun, Zhaochun Ren, Zhaopeng Tu, and Maarten de Rijke. 2020 · 2020
Cited alongside, same era.
Deconstruct to Reconstruct a Configurable Evaluation Metric for Open-Domain Dialogue Systems
Vitou Phy, Yang Zhao, and Akiko Aizawa. 2020 · 2020
Cited alongside, same era.
BoB: BERT Over BERT for Training Persona-based Dialogue Models from Limited Personalized Data
Haoyu Song, Yan Wang, Kaiyan Zhang, Wei-Nan Zhang, and Ting Liu. 2021 · 2021
Later among the works it cites.
Underreporting of errors in NLG output, and what to do about it
Emiel van Miltenburg, Miruna Clinciu, Ondřej Dušek, Dimitra Gkatzia, Stephanie Inglis, Leo Leppänen, Saad Mahamood, Emma Manning, Stephanie Schoch, Craig Thomson, and Luou Wen. 2021 · 2021
Later among the works it cites.
Blender Bot 2.0: An open source chatbot that builds long-term memory and searches the internet
Jason Weston and Kurt Shuster. 2021 · 2021
Later among the works it cites.
Bot-Adversarial Dialogue for Safe Conversational Agents
Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan. 2021 · 2021
Later among the works it cites.
CoLV: A Collaborative Latent Variable Model for Knowledge-Grounded Dialogue Generation
Haolan Zhan, Lei Shen, Hongshen Chen, and Hainan Zhang. 2021 · 2021
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Annotating Errors and Emotions in Human-Chatbot Interactions in Italian
Manuela Sanguinetti, Alessandro Mazzei, Viviana Patti, Marco Scalerandi, Dario Mana, and Rossana Simeoni. 2020 · 2020
Cited alongside, same era.
Item Response Theory for Efficient Human Evaluation of Chatbots
João Sedoc and Lyle Ungar. 2020 · 2020
Cited alongside, same era.
Generate, Delete and Rewrite: A Three-Stage Framework for Improving Persona Consistency of Dialogue Generation
Haoyu Song, Yan Wang, Wei-Nan Zhang, Xiaojiang Liu, and Ting Liu. 2020 · 2020
Cited alongside, same era.
Knowledge-Grounded Response Generation with Deep Attentional Latent-Variable Model
Hao-Tong Ye, Kai-Lin Lo, Shang-Yu Su, and Yun-Nung Chen. 2020 · 2020
Cited alongside, same era.
DIALOGPT : Large-Scale Generative Pre-training for Conversational Response Generation
Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan. 2020 · 2020
Cited alongside, same era.
Integrated taxonomy of errors in chat-oriented dialogue systems
Ryuichiro Higashinaka, Masahiro Araki, Hiroshi Tsukahara, and Masahiro Mizukami. 2021 · 2021
Cited alongside, same era.
Q2: Evaluating Factual Consistency in Knowledge-Grounded Dialogues via Question Generation and Question Answering
Or Honovich, Leshem Choshen, Roee Aharoni, Ella Neeman, Idan Szpektor, and Omri Abend. 2021 · 2021
Cited alongside, same era.
CARE: Commonsense-Aware Emotional Response Generation with Latent Concepts
Peixiang Zhong, Di Wang, Pengfei Li, Chen Zhang, Hao Wang, and Chunyan Miao. 2021 · 2021
Later among the works it cites.
Commonsense-Focused Dialogues for Response Generation: An Empirical Study
Pei Zhou, Karthik Gopalakrishnan, Behnam Hedayatnia, Seokhwan Kim, Jay Pujara, Xiang Ren, Yang Liu, and Dilek Hakkani-Tur. 2021 · 2021
Later among the works it cites.
Measuring and Mitigating Bias in AI-Chatbots
Hedin Beattie, Lanier Watkins, William H. Robinson, Aviel Rubin, and Shari Watkins. 2022 · 2022
Closest in time.
Probing the Robustness of Trained Metrics for Conversational Dialogue Systems
Jan Deriu, Don Tuggener, Pius Von Däniken, and Mark Cieliebak. 2022 · 2022
Closest in time.
Safetykit: First aid for measuring safety in open-domain conversational systems
Emily Dinan, Gavin Abercrombie, A Bergman, Shannon L Spruit, Dirk Hovy, Y-Lan Boureau, and Verena Rieser. 2022 · 2022
Closest in time.
DialFact: A Benchmark for Fact-Checking in Dialogue
Prakhar Gupta, Chien-Sheng Wu, Wenhao Liu, and Caiming Xiong. 2022 · 2022
Closest in time.
Achieving Reliable Human Assessment of Open-Domain Dialogue Systems
Tianbo Ji, Yvette Graham, Gareth Jones, Chenyang Lyu, and Qun Liu. 2022 · 2022
Closest in time.
Internet-Augmented Dialogue Generation
Mojtaba Komeili, Kurt Shuster, and Jason Weston. 2022 · 2022
Closest in time.
CEM: Commonsense-Aware Empathetic Response Generation
Sahand Sabour, Chujie Zheng, and Minlie Huang. 2022 · 2022
Closest in time.
Human Evaluation of Conversations is an Open Problem: comparing the sensitivity of various methods for evaluating dialogue agents
Eric Smith, Orion Hsu, Rebecca Qian, Stephen Roller, Y-Lan Boureau, and Jason Weston. 2022 · 2022
Closest in time.
On the Safety of Conversational Models: Taxonomy, Dataset, and Benchmark
Hao Sun, Guangxuan Xu, Jiawen Deng, Jiale Cheng, Chujie Zheng, Hao Zhou, Nanyun Peng, Xiaoyan Zhu, and Minlie Huang. 2022 · 2022
Closest in time.
Beyond Goldfish Memory: Long-Term Open-Domain Conversation
Jing Xu, Arthur Szlam, and Jason Weston. 2022 · 2022
Closest in time.
Think Before You Speak: Explicitly Generating Implicit Commonsense Knowledge for Response Generation
Pei Zhou, Karthik Gopalakrishnan, Behnam Hedayatnia, Seokhwan Kim, Jay Pujara, Xiang Ren, Yang Liu, and Dilek Hakkani-Tur. 2022 · 2022
Closest in time.
Can You Put it All Together: Evaluating Conversational Agents’ Ability to Blend Skills
Eric Michael Smith, Mary Williamson, Kurt Shuster, Jason Weston, and Y-Lan Boureau. 2020 · 2030
Closest in time.