Fetching the paper…
Reading the bibliography…
Open-domain dialog systems have a user-centric goal: to provide humans with an engaging conversation experience.
The second conversational intelligence challenge (convai2)
Emily Dinan, Varvara Logacheva, Valentin Malykh, Alexander Miller, Kurt Shuster, Jack Urbanek, Douwe Kiela, Arthur Szlam, Iulian Serban, Ryan Lowe, et al. 2019 · 1902
Earlier work this paper cites.
Sanghyun Yi, Rahul Goel, Chandra Khatri, Tagyoung Chung, Behnam Hedayatnia, Anu Venkatesh, Raefer Gabriel, and Dilek Hakkani-Tür. 2019 · 1904
Earlier work this paper cites.
Investigating evaluation of open-domain dialogue systems with human generated multiple references
Prakhar Gupta, Shikib Mehri, Tiancheng Zhao, Amy Pavel, Maxine Eskénazi, and Jeffrey P. Bigham. 2019 · 1907
Earlier work this paper cites.
ACUTE-EVAL: improved dialogue evaluation with optimized questions and multi-turn comparisons
Margaret Li, Jason Weston, and Stephen Roller. 2019 · 1909
Earlier work this paper cites.
Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit
Jacob Cohen. 1968 · 1968
Earlier work this paper cites.
On the uniqueness of the shapley value
Pradeep Dubey. 1975 · 1975
Earlier work this paper cites.
Bargaining foundations of shapley value
Faruk Gul. 1989 · 1989
Earlier work this paper cites.
Modern information retrieval , volume 463
Ricardo Baeza-Yates, Berthier Ribeiro-Neto, et al. 1999 · 1999
Earlier work this paper cites.
Towards a human-like open-domain chatbot
Daniel Adiwardana, Minh-Thang Luong, David R. So, Jamie Hall, Noah Fiedel, Romal Thoppilan, Zi Yang, Apoorv Kulshreshtha, Gaurav Nemade, Yifeng Lu, and Quoc V. Le. 2020 · 2001
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Recipes for building an open-domain chatbot
Stephen Roller, Emily Dinan, Naman Goyal, Da Ju, Mary Williamson, Yinhan Liu, Jing Xu, Myle Ott, Kurt Shuster, Eric Michael Smith, Y-Lan Boureau, and Jason Weston. 2020 · 2004
Earlier work this paper cites.
Detecting user engagement in everyday conversations
Chen Yu, Paul M. Aoki, and Allison Woodruff. 2004 · 2004
Earlier work this paper cites.
METEOR: an automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Learning to extract relations from the web using minimal supervision
Razvan C. Bunescu and Raymond J. Mooney. 2007 · 2007
Earlier work this paper cites.
Distant supervision for relation extraction without labeled data
Mike Mintz, Steven Bills, Rion Snow, and Daniel Jurafsky. 2009 · 2009
Earlier work this paper cites.
Gunrock 2.0: A user adaptive social conversational system
Kaihui Liang, Austin Chau, Yu Li, Xueyuan Lu, Dian Yu, Mingyang Zhou, Ishan Jain, Sam Davidson, Josh Arnold, Minh Nguyen, et al. 2020a · 2011
Earlier work this paper cites.
Weixin Liang, Feiyang Niu, Aishwarya N. Reganti, Govind Thattai, and Gökhan Tür. 2020b · 2011
Cited alongside, same era.
A comparison of greedy and optimal assessment of natural language student input using word-to-word similarity metrics
Vasile Rus and Mihai C. Lintean. 2012 · 2012
Cited alongside, same era.
Bootstrapping dialog systems with word embeddings
Gabriel Forgues, Joelle Pineau, Jean-Marie Larchevêque, and Réal Tremblay. 2014 · 2014
Cited alongside, same era.
How NOT to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation
Chia-Wei Liu, Ryan Lowe, Iulian Serban, Michael Noseworthy, Laurent Charlin, and Joelle Pineau. 2016 · 2016
Cited alongside, same era.
Data programming: Creating large training sets, quickly
Alexander J. Ratner, Christopher De Sa, Sen Wu, Daniel Selsam, and Christopher Ré. 2016 · 2016
RUBER: an unsupervised method for automatic evaluation of open-domain dialog systems
Chongyang Tao, Lili Mou, Dongyan Zhao, and Rui Yan. 2018 · 2018
Later among the works it cites.
On evaluating and comparing conversational agents
Anu Venkatesh, Chandra Khatri, Ashwin Ram, Fenfei Guo, Raefer Gabriel, Ashish Nagar, Rohit Prasad, Ming Cheng, Behnam Hedayatnia, Angeliki Metallinou, Rahul Goel, Shaohua Yang, and Anirudh Raju. 2018 · 2018
Later among the works it cites.
Offline and online satisfaction prediction in open-domain conversational systems
Jason Ingyu Choi, Ali Ahmadvand, and Eugene Agichtein. 2019 · 2019
Later among the works it cites.
BERT: pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Data shapley: Equitable valuation of data for machine learning
Amirata Ghorbani and James Y. Zou. 2019 · 2019
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
A wizard-of-oz study on A non-task-oriented dialog systems that reacts to user engagement
Zhou Yu, Leah Nicolich-Henkin, Alan W. Black, and Alexander I. Rudnicky. 2016 · 2016
Cited alongside, same era.
Understanding black-box predictions via influence functions
Pang Wei Koh and Percy Liang. 2017 · 2017
Cited alongside, same era.
Towards an automatic turing test: Learning to evaluate dialogue responses
Ryan Lowe, Michael Noseworthy, Iulian Vlad Serban, Nicolas Angelard-Gontier, Yoshua Bengio, and Joelle Pineau. 2017 · 2017
Cited alongside, same era.
Why we need new evaluation metrics for NLG
Jekaterina Novikova, Ondrej Dusek, Amanda Cercas Curry, and Verena Rieser. 2017 · 2017
Cited alongside, same era.
Inferring generative model structure with static analysis
Paroma Varma, Bryan D. He, Payal Bajaj, Nishith Khandwala, Imon Banerjee, Daniel L. Rubin, and Christopher Ré. 2017 · 2017
Cited alongside, same era.
An open-source dialog system with real-time engagement tracking for job interview training applications
Zhou Yu, Vikram Ramanarayanan, Patrick L. Lange, and David Suendermann-Oeft. 2017 · 2017
Cited alongside, same era.
Gunrock: Building a human-like social bot by leveraging large scale real user data
Chun-Yen Chen, Dian Yu, Weiming Wen, Yi Mang Yang, Jiaping Zhang, Mingyang Zhou, Kevin Jesse, Austin Chau, Antara Bhowmick, Shreenath Iyer, et al. 2018 · 2018
Cited alongside, same era.
Learning from dialogue after deployment: Feed yourself, chatbot!
Braden Hancock, Antoine Bordes, Pierre-Emmanuel Mazaré, and Jason Weston. 2019 · 2019
Later among the works it cites.
Scene graph prediction with limited labels
Ranjay Krishna, Vincent S. Chen, Paroma Varma, Michael Bernstein, Christopher Ré, and Fei-Fei Li. 2019 · 2019
Later among the works it cites.
Bootstrapping conversational agents with weak supervision
Neil Mallinar, Abhishek Shah, Rajendra Ugrani, Ayush Gupta, Manikandan Gurusankar, Tin Kam Ho, Q. Vera Liao, Yunfeng Zhang, Rachel K. E. Bellamy, Robert Yates, Chris Desmarais, and Blake McGregor. 2019 · 2019
Later among the works it cites.
Predictive engagement: An efficient metric for automatic evaluation of open-domain dialogue systems
Sarik Ghazarian, Ralph M. Weischedel, Aram Galstyan, and Nanyun Peng. 2020 · 2020
Later among the works it cites.
Unsupervised evaluation of interactive dialog with dialogpt
Shikib Mehri and Maxine Eskénazi. 2020 · 2020
Later among the works it cites.
Estimating training data influence by tracing gradient descent
Garima Pruthi, Frederick Liu, Satyen Kale, and Mukund Sundararajan. 2020 · 2020
Later among the works it cites.
Snorkel: rapid training data creation with weak supervision
Alexander Ratner, Stephen H. Bach, Henry R. Ehrenberg, Jason A. Fries, Sen Wu, and Christopher Ré. 2020 · 2020
Later among the works it cites.
Privacy in crowdsourcing: a review of the threats and challenges
Huichuan Xia and Brian McKernan. 2020 · 2020
Later among the works it cites.
GraghVQA: Language-guided graph neural networks for graph-based visual question answering
Weixin Liang, Yanhao Jiang, and Zixuan Liu. 2021 · 2021
Closest in time.
Neural group testing to accelerate deep learning
Weixin Liang and James Zou. 2021 · 2021
Closest in time.
Midas: A dialog act annotation scheme for open domain human machine spoken conversations
Dian Yu and Zhou Yu. 2021 · 2021
Closest in time.