Fetching the paper…
Reading the bibliography…
At the heart of improving conversational AI is the open problem of how to evaluate conversations.
The second conversational intelligence challenge (convai2)
Emily Dinan, Varvara Logacheva, Valentin Malykh, Alexander Miller, Kurt Shuster, Jack Urbanek, Douwe Kiela, Arthur Szlam, Iulian Serban, Ryan Lowe, et al. 2019b · 1902
Earlier work this paper cites.
What makes a good conversation? how controllable attributes affect human judgments
Abigail See, Stephen Roller, Douwe Kiela, and Jason Weston. 2019b · 1902
Earlier work this paper cites.
Unifying human and statistical evaluation for natural language generation
Tatsunori B Hashimoto, Hugh Zhang, and Percy Liang. 2019 · 1904
Earlier work this paper cites.
Eli5: Long form question answering
Angela Fan, Yacine Jernite, Ethan Perez, David Grangier, Jason Weston, and Michael Auli. 2019 · 1907
Earlier work this paper cites.
Investigating evaluation of open-domain dialogue systems with human generated multiple references
Prakhar Gupta, Shikib Mehri, Tiancheng Zhao, Amy Pavel, Maxine Eskenazi, and Jeffrey P Bigham. 2019 · 1907
Earlier work this paper cites.
Build it break it fix it for dialogue safety: Robustness from adversarial human attack
Emily Dinan, Samuel Humeau, Bharath Chintagunta, and Jason Weston. 2019a · 1908
Earlier work this paper cites.
Acute-eval: Improved dialogue evaluation with optimized questions and multi-turn comparisons
Margaret Li, Jason Weston, and Stephen Roller. 2019 · 1909
Earlier work this paper cites.
Forming impressions of personality
SE Asch. 1946 · 1946
Earlier work this paper cites.
Computing machinery and intelligence
Alan M Turing and J Haugeland. 1950 · 1950
Earlier work this paper cites.
Recency and primacy in persuasion as a function of the timing of speeches and measurements
Norman Miller and Donald T Campbell. 1959 · 1959
Earlier work this paper cites.
The serial position effect of free recall
Bennet B Murdock Jr. 1962 · 1962
Earlier work this paper cites.
Primacy effects in personality impression formation using a generalized order effect paradigm
Norman H Anderson. 1965 · 1965
Earlier work this paper cites.
Short-term temporal changes in free recall
Leo Postman and Laura W Phillips. 1965 · 1965
Earlier work this paper cites.
Availability and interference in predictive judgment
Stephen J Hoch. 1984 · 1984
Earlier work this paper cites.
Towards a human-like open-domain chatbot
Daniel Adiwardana, Minh-Thang Luong, David R So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, et al. 2020 · 2001
Earlier work this paper cites.
Jason Baumgartner, Savvas Zannettou, Brian Keegan, Megan Squire, and Jeremy Blackburn. 2020 · 2001
Earlier work this paper cites.
Beyond user self-reported likert scale ratings: A comparison model for automatic dialog evaluation
Weixin Liang, James Zou, and Zhou Yu. 2020 · 2005
Earlier work this paper cites.
Absolute identification by relative judgment
Neil Stewart, Gordon DA Brown, and Nick Chater. 2005 · 2005
Earlier work this paper cites.
Open-domain conversational agents: Current progress, open problems, and future directions
Stephen Roller, Y-Lan Boureau, Jason Weston, Antoine Bordes, Emily Dinan, Angela Fan, David Gunning, Da Ju, Margaret Li, Spencer Poff, et al. 2020 · 2006
Earlier work this paper cites.
A re-examination of machine learning approaches for sentence-level mt evaluation
Joshua Albrecht and Rebecca Hwa. 2007 · 2007
Cited alongside, same era.
Deploying lifelong open-domain dialogue learning
Kurt Shuster, Jack Urbanek, Emily Dinan, Arthur Szlam, and Jason Weston. 2020 · 2008
Cited alongside, same era.
Spot the bot: A robust and efficient framework for the evaluation of conversational dialogue systems
Jan Deriu, Don Tuggener, Pius von Däniken, Jon Ander Campos, Alvaro Rodrigo, Thiziri Belkacem, Aitor Soroa, Eneko Agirre, and Mark Cieliebak. 2020 · 2010
Cited alongside, same era.
An evaluation protocol for generative conversational systems
Seolhwa Lee, Heuiseok Lim, and João Sedoc. 2020 · 2010
Cited alongside, same era.
Personalizing dialogue agents: I have a dog, do you have pets too?
Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston. 2018 · 2018
Later among the works it cites.
Evaluating question answering evaluation
Anthony Chen, Gabriel Stanovsky, Sameer Singh, and Matt Gardner. 2019 · 2019
Later among the works it cites.
Approximating interactive human evaluation with self-play for open-domain dialog systems
Asma Ghandeharioun, Judy Hanwen Shen, Natasha Jaques, Craig Ferguson, Noah Jones, Àgata Lapedriza, and Rosalind W. Picard. 2019 · 2019
Later among the works it cites.
Importance of search and evaluation strategies in neural dialogue modeling
Ilia Kulikov, Alexander Miller, Kyunghyun Cho, and Jason Weston. 2019 · 2019
Later among the works it cites.
Towards empathetic open-domain conversation models: A new benchmark and dataset
Hannah Rashkin, Eric Michael Smith, Margaret Li, and Y-Lan Boureau. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Jing Xu, Da Ju, Margaret Li, Y-Lan Boureau, Jason Weston, and Emily Dinan. 2020 · 2010
Cited alongside, same era.
Metrics and evaluation of spoken dialogue systems
Helen Hastie. 2012 · 2012
Cited alongside, same era.
The second dialog state tracking challenge
Matthew Henderson, Blaise Thomson, and Jason D Williams. 2014 · 2014
Cited alongside, same era.
Oriol Vinyals and Quoc Le. 2015 · 2015
Cited alongside, same era.
A persona-based neural conversation model
Jiwei Li, Michel Galley, Chris Brockett, Georgios P Spithourakis, Jianfeng Gao, and Bill Dolan. 2016 · 2016
Cited alongside, same era.
Chia-Wei Liu, Ryan Lowe, Iulian V Serban, Michael Noseworthy, Laurent Charlin, and Joelle Pineau. 2016 · 2016
Cited alongside, same era.
The dialog state tracking challenge series: A review
Jason D Williams, Antoine Raux, and Matthew Henderson. 2016 · 2016
Cited alongside, same era.
Conversational AI: The science behind the Adlexa Prize
Ram Ashwin, Prasad Rohit, Khatri Chandra, Venkatesh Anu, Gabriel Raefer, Liu Qing, Nunn Jeff, Hedayatnia Behnam, Cheng Ming, Nagar Ashish, King Eric, Bland Kate, Wartick Amanda, Pan Yi, Song Han, Jayadevan Sk, Hwang Gene, and Pettigrue Art. 2017 · 2017
Cited alongside, same era.
What makes a good conversation? how controllable attributes affect human judgments
Abigail See, Stephen Roller, Douwe Kiela, and Jason Weston. 2019a · 2019
Later among the works it cites.
Response generation by context-aware prototype editing
Yu Wu, Furu Wei, Shaohan Huang, Yunli Wang, Zhoujun Li, and Ming Zhou. 2019 · 2019
Later among the works it cites.
Further advances in open domain dialog systems in the third alexa prize socialbot grand challenge
Raefer Gabriel, Yang Liu, Anna Gottardi, Mihail Eric, Anju Khatri, Anjali Chadha, Qinlang Chen, Behnam Hedayatnia, Pankaj Rajan, Ali Binici, et al. 2020 · 2020
Later among the works it cites.
Challenges in building intelligent open-domain dialog systems
Minlie Huang, Xiaoyan Zhu, and Jianfeng Gao. 2020 · 2020
Later among the works it cites.
Neural text generation with unlikelihood training
Sean Welleck, Ilia Kulikov, Stephen Roller, Emily Dinan, Kyunghyun Cho, and Jason Weston. 2020 · 2020
Later among the works it cites.
Reason first, then respond: Modular generation for knowledge-infused dialogue
Leonard Adolphs, Kurt Shuster, Jack Urbanek, Arthur Szlam, and Jason Weston. 2021 · 2021
Later among the works it cites.
Choose your own adventure: Paired suggestions in collaborative writing for evaluating story generation models
Elizabeth Clark and Noah A Smith. 2021 · 2021
Later among the works it cites.
Survey on evaluation methods for dialogue systems
Jan Deriu, Alvaro Rodrigo, Arantxa Otegi, Guillermo Echegoyen, Sophie Rosset, Eneko Agirre, and Mark Cieliebak. 2021 · 2021
Later among the works it cites.
A survey of nlp-related crowdsourcing hits: what works and what does not
Jessica Huynh, Jeffrey Bigham, and Maxine Eskenazi. 2021 · 2021
Later among the works it cites.
Internet-augmented dialogue generation
Mojtaba Komeili, Kurt Shuster, and Jason Weston. 2021 · 2021
Later among the works it cites.
Recipes for building an open-domain chatbot
Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Eric Michael Smith, Y-Lan Boureau, et al. 2021 · 2021
Later among the works it cites.
Beyond goldfish memory: Long-term open-domain conversation
Jing Xu, Arthur Szlam, and Jason Weston. 2021 · 2021
Later among the works it cites.
A comprehensive assessment of dialog evaluation metrics
Yi-Ting Yeh, Maxine Eskenazi, and Shikib Mehri. 2021 · 2021
Later among the works it cites.