Fetching the paper…
Reading the bibliography…
There is a multitude of novel generative models for open-domain conversational systems; however, there is no systematic evaluation of different systems.
The second conversational intelligence challenge (convai2)
Emily Dinan, Varvara Logacheva, Valentin Malykh, Alexander Miller, Kurt Shuster, Jack Urbanek, Douwe Kiela, Arthur Szlam, Iulian Serban, Ryan Lowe, et al. 2019 · 1902
Earlier work this paper cites.
Deep learning based chatbot models
Richard Csaky. 2019 · 1908
Earlier work this paper cites.
Acute-eval: Improved dialogue evaluation with optimized questions and multi-turn comparisons
Margaret Li, Jason Weston, and Stephen Roller. 2019 · 1909
Earlier work this paper cites.
The dialogue dodecathlon: Open-domain knowledge and image grounded conversational agents
Kurt Shuster, Da Ju, Stephen Roller, Emily Dinan, Y-Lan Boureau, and Jason Weston. 2019 · 1911
Earlier work this paper cites.
Dialogpt: Large-scale generative pre-training for conversational response generation
Yizhe Zhang, Siqi Sun, Michel Galley, Yen-Chun Chen, Chris Brockett, Xiang Gao, Jianfeng Gao, Jingjing Liu, and Bill Dolan. 2019 · 1911
Earlier work this paper cites.
Rank analysis of incomplete block designs: I. the method of paired comparisons
Ralph Allan Bradley and Milton E Terry. 1952 · 1952
Earlier work this paper cites.
Correction of item-total correlations in item analysis
Sten Henrysson. 1963 · 1963
Earlier work this paper cites.
Measuring nominal scale agreement among many raters
Joseph L Fleiss. 1971 · 1971
Earlier work this paper cites.
Towards a human-like open-domain chatbot
Daniel Adiwardana, Minh-Thang Luong, David R So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, et al. 2020 · 2001
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Recipes for building an open-domain chatbot
Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Kurt Shuster, Eric M Smith, et al. 2020 · 2004
Earlier work this paper cites.
Can you put it all together: Evaluating conversational agents’ ability to blend skills
Eric Michael Smith, Mary Williamson, Kurt Shuster, Jason Weston, and Y-Lan Boureau. 2020 · 2004
Earlier work this paper cites.
Language models are few-shot learners
Tom B Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. 2020 · 2005
Earlier work this paper cites.
Trueskill™: a bayesian skill rating system
Ralf Herbrich, Tom Minka, and Thore Graepel. 2007 · 2007
Earlier work this paper cites.
Chameleons in imagined conversations: A new approach to understanding coordination of linguistic style in dialogs
Cristian Danescu-Niculescu-Mizil and Lillian Lee. 2011 · 2011
Earlier work this paper cites.
Toward learning and evaluation of dialogue policies with text examples
David DeVault, Anton Leuski, and Kenji Sagae. 2011 · 2011
Earlier work this paper cites.
An empirical investigation of statistical significance in nlp
Taylor Berg-Kirkpatrick, David Burkett, and Dan Klein. 2012 · 2012
Earlier work this paper cites.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D Manning. 2014 · 2014
Cited alongside, same era.
Efficient elicitation of annotations for human evaluation of machine translation
Keisuke Sakaguchi, Matt Post, and Benjamin Van Durme. 2014 · 2014
Cited alongside, same era.
Oriol Vinyals and Quoc Le. 2015 · 2015
Cited alongside, same era.
The dialogue breakdown detection challenge: Task description, datasets, and evaluation metrics
Ryuichiro Higashinaka, Kotaro Funakoshi, Yuka Kobayashi, and Michimasa Inaba. 2016 · 2016
Cited alongside, same era.
A diversity-promoting objective function for neural conversation models
Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. 2016a · 2016
Cited alongside, same era.
The first conversational intelligence challenge
Mikhail Burtsev, Varvara Logacheva, Valentin Malykh, Iulian Vlad Serban, Ryan Lowe, Shrimai Prabhumoye, Alan W Black, Alexander Rudnicky, and Yoshua Bengio. 2018 · 2018
Later among the works it cites.
Wizard of wikipedia: Knowledge-powered conversational agents
Emily Dinan, Stephen Roller, Kurt Shuster, Angela Fan, Michael Auli, and Jason Weston. 2018 · 2018
Later among the works it cites.
The hitchhiker’s guide to testing statistical significance in natural language processing
Rotem Dror, Gili Baumer, Segev Shlomov, and Roi Reichart. 2018 · 2018
Later among the works it cites.
A dataset of topic-oriented human-to-chatbot dialogues
Varvara Logacheva, Mikhail Burtsev, Valentin Malykh, Vadim Poluliakh, Alexander Rudnicky, Iulian Serban, Ryan Lowe, Shrimai Prabhumoye, Alan W Black, and Yoshua Bengio. 2018 · 2018
Later among the works it cites.
RankME: Reliable human ratings for natural language generation
Jekaterina Novikova, Ondrej Dušek, and Verena Rieser. 2018 · 2018
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
How NOT to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation
Chia-Wei Liu, Ryan Lowe, Iulian Serban, Mike Noseworthy, Laurent Charlin, and Joelle Pineau. 2016 · 2016
Cited alongside, same era.
Overview of dialogue breakdown detection challenge 3
Ryuichiro Higashinaka, Kotaro Funakoshi, Michimasa Inaba, Yuiko Tsunomori, Tetsuro Takahashi, and Nobuhiro Kaji. 2017 · 2017
Cited alongside, same era.
OpenNMT: Open-source toolkit for neural machine translation
Guillaume Klein, Yoon Kim, Yuntian Deng, Jean Senellart, and Alexander M. Rush. 2017 · 2017
Cited alongside, same era.
Adversarial learning for neural dialogue generation
Jiwei Li, Will Monroe, Tianlin Shi, Sébastien Jean, Alan Ritter, and Dan Jurafsky. 2017a · 2017
Cited alongside, same era.
Towards an automatic turing test: Learning to evaluate dialogue responses
Ryan Lowe, Michael Noseworthy, Iulian V Serban, Nicolas Angelard-Gontier, Yoshua Bengio, and Joelle Pineau. 2017 · 2017
Cited alongside, same era.
ParlAI: A dialog research software platform
Alexander Miller, Will Feng, Dhruv Batra, Antoine Bordes, Adam Fisch, Jiasen Lu, Devi Parikh, and Jason Weston. 2017 · 2017
Cited alongside, same era.
Why we need new evaluation metrics for NLG
Jekaterina Novikova, Ondřej Dušek, Amanda Cercas Curry, and Verena Rieser. 2017 · 2017
Cited alongside, same era.
On Evaluating and Comparing Conversational Agents
Anu Venkatesh, Chandra Khatri, Ashwin Ram, Fenfei Guo, Raefer Gabriel, Ashish Nagar, Rohit Prasad, Ming Cheng, Behnam Hedayatnia, Angeliki Metallinou, Rahul Goel, Shaohua Yang, and Anirudh Raju. 2018 · 2018
Later among the works it cites.
Personalizing dialogue agents: I have a dog, do you have pets too?
Saizheng Zhang, Emily Dinan, Jack Urbanek, Arthur Szlam, Douwe Kiela, and Jason Weston. 2018 · 2018
Later among the works it cites.
Agreement is overrated: A plea for correlation to assess human evaluation reliability
Jacopo Amidei, Paul Piwek, and Alistair Willis. 2019 · 2019
Later among the works it cites.
Improving neural conversational models with entropy-based data filtering
Richárd Csáky, Patrik Purgai, and Gábor Recski. 2019 · 2019
Later among the works it cites.
Neural Approaches to Conversational AI: Question Answering, Task-oriented Dialogues and Social Chatbots
Jianfeng Gao, Michel Galley, and Lihong Li. 2019 · 2019
Later among the works it cites.
Language models are unsupervised multitask learners
Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019 · 2019
Later among the works it cites.
Towards empathetic open-domain conversation models: A new benchmark and dataset
Hannah Rashkin, Eric Michael Smith, Margaret Li, and Y-Lan Boureau. 2019 · 2019
Later among the works it cites.
ChatEval: A tool for chatbot evaluation
João Sedoc, Daphne Ippolito, Arun Kirubarajan, Jai Thirani, Lyle Ungar, and Chris Callison-Burch. 2019 · 2019
Later among the works it cites.
What makes a good conversation? how controllable attributes affect human judgments
Abigail See, Stephen Roller, Douwe Kiela, and Jason Weston. 2019 · 2019
Later among the works it cites.
The pushshift reddit dataset
Jason Baumgartner, Savvas Zannettou, Brian Keegan, Megan Squire, and Jeremy Blackburn. 2020 · 2020
Closest in time.
Overview of the seventh dialog system technology challenge: Dstc7
Luis Fernando D’Haro, Koichiro Yoshino, Chiori Hori, Tim K Marks, Lazaros Polymenakos, Jonathan K Kummerfeld, Michel Galley, and Xiang Gao. 2020 · 2020
Closest in time.