Fetching the paper…
Reading the bibliography…
The aim of this paper is to mitigate the shortcomings of automatic evaluation of open-domain dialog systems through multi-reference evaluation.
Jointly optimizing diversity and relevance in neural response generation
Xiang Gao, Sungjin Lee, Yizhe Zhang, Chris Brockett, Michel Galley, Jianfeng Gao, and Bill Dolan. 2019 · 1902
Earlier work this paper cites.
Re-evaluating adem: A deeper look at scoring dialogue responses
Ananya B Sai, Mithun Das Gupta, Mitesh M Khapra, and Mukundhan Srinivasan. 2019 · 1902
Earlier work this paper cites.
Unifying human and statistical evaluation for natural language generation
Tatsunori B Hashimoto, Hugh Zhang, and Percy Liang. 2019 · 1904
Earlier work this paper cites.
Survey on evaluation methods for dialogue systems
Jan Deriu, Alvaro Rodrigo, Arantxa Otegi, Guillermo Echegoyen, Sophie Rosset, Eneko Agirre, and Mark Cieliebak. 2019 · 1905
Earlier work this paper cites.
Weighted kappa: Nominal scale agreement provision for scaled disagreement or partial credit
Jacob Cohen. 1968 · 1968
Earlier work this paper cites.
Bleu: A method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with high levels of correlation with human judgments
Alon Lavie and Abhaya Agarwal. 2007 · 2007
Earlier work this paper cites.
Data-driven response generation in social media
Alan Ritter, Colin Cherry, and William B Dolan. 2011 · 2011
Earlier work this paper cites.
A comparison of greedy and optimal assessment of natural language student input using word-to-word similarity metrics
Vasile Rus and Mihai Lintean. 2012 · 2012
Earlier work this paper cites.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2014 · 2014
Earlier work this paper cites.
Bootstrapping dialog systems with word embeddings
Gabriel Forgues, Joelle Pineau, Jean-Marie Larchevêque, and Réal Tremblay. 2014 · 2014
Earlier work this paper cites.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014 · 2014
Cited alongside, same era.
deltaBLEU: A discriminative metric for generation tasks with intrinsically diverse targets
Michel Galley, Chris Brockett, Alessandro Sordoni, Yangfeng Ji, Michael Auli, Chris Quirk, Margaret Mitchell, Jianfeng Gao, and Bill Dolan. 2015 · 2015
Cited alongside, same era.
Skip-thought vectors
Ryan Kiros, Yukun Zhu, Ruslan R Salakhutdinov, Richard Zemel, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Cited alongside, same era.
Truly exploring multiple references for machine translation evaluation
Ying Qin and Lucia Specia. 2015 · 2015
Cited alongside, same era.
A neural network approach to context-sensitive generation of conversational responses
Alessandro Sordoni, Michel Galley, Michael Auli, Chris Brockett, Yangfeng Ji, Margaret Mitchell, Jian-Yun Nie, Jianfeng Gao, and William B. Dolan. 2015 · 2015
Cited alongside, same era.
Towards an automatic Turing test: Learning to evaluate dialogue responses
Ryan Lowe, Michael Noseworthy, Iulian Vlad Serban, Nicolas Angelard-Gontier, Yoshua Bengio, and Joelle Pineau. 2017 · 2017
Later among the works it cites.
A hierarchical latent variable encoder-decoder model for generating dialogues
Iulian Vlad Serban, Alessandro Sordoni, Ryan Lowe, Laurent Charlin, Joelle Pineau, Aaron Courville, and Yoshua Bengio. 2017 · 2017
Later among the works it cites.
Shikhar Sharma, Layla El Asri, Hannes Schulz, and Jeremie Zumer. 2017 · 2017
Later among the works it cites.
A knowledge-grounded neural conversation model
Marjan Ghazvininejad, Chris Brockett, Ming-Wei Chang, Bill Dolan, Jianfeng Gao, Wen-tau Yih, and Michel Galley. 2018 · 2018
Later among the works it cites.
Importance of a search strategy in neural dialogue modelling
Ilya Kulikov, Alexander H Miller, Kyunghyun Cho, and Jason Weston. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Oriol Vinyals and Quoc Le. 2015 · 2015
Cited alongside, same era.
Towards universal paraphrastic sentence embeddings
John Wieting, Mohit Bansal, Kevin Gimpel, and Karen Livescu. 2015 · 2015
Cited alongside, same era.
Generating sentences from a continuous space
Samuel R Bowman, Luke Vilnis, Oriol Vinyals, Andrew Dai, Rafal Jozefowicz, and Samy Bengio. 2016 · 2016
Cited alongside, same era.
A diversity-promoting objective function for neural conversation models
Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. 2016a · 2016
Cited alongside, same era.
How NOT to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation
Chia-Wei Liu, Ryan Lowe, Iulian Serban, Mike Noseworthy, Laurent Charlin, and Joelle Pineau. 2016 · 2016
Cited alongside, same era.
DailyDialog: A manually labelled multi-turn dialogue dataset
Yanran Li, Hui Su, Xiaoyu Shen, Wenjie Li, Ziqiang Cao, and Shuzi Niu. 2017 · 2017
Cited alongside, same era.
A persona-based neural conversation model
Jiwei Li, Michel Galley, Chris Brockett, Georgios Spithourakis, Jianfeng Gao, and Bill Dolan. 2016b
Cited in the paper.
Later among the works it cites.
Learning general purpose distributed sentence representations via large scale multi-task learning
Sandeep Subramanian, Adam Trischler, Yoshua Bengio, and Christopher J Pal. 2018 · 2018
Later among the works it cites.
Ruber: An unsupervised method for automatic evaluation of open-domain dialog systems
Chongyang Tao, Lili Mou, Dongyan Zhao, and Rui Yan. 2018 · 2018
Later among the works it cites.
Generating informative and diverse conversational responses via adversarial information maximization
Yizhe Zhang, Michel Galley, Jianfeng Gao, Zhe Gan, Xiujun Li, Chris Brockett, and Bill Dolan. 2018 · 2018
Later among the works it cites.
Texygen: A benchmarking platform for text generation models
Yaoming Zhu, Sidi Lu, Lei Zheng, Jiaxian Guo, Weinan Zhang, Jun Wang, and Yong Yu. 2018 · 2018
Later among the works it cites.
Automatic Evaluation of Chat-Oriented Dialogue Systems Using Large-Scale Multi-references , pages 15–25. Springer International Publishing, Cham
Hiroaki Sugiyama, Toyomi Meguro, and Ryuichiro Higashinaka. 2019 · 2019
Closest in time.
LSDSCC: a large scale domain-specific conversational corpus for response generation with diversity oriented evaluation metrics
Zhen Xu, Nan Jiang, Bingquan Liu, Wenge Rong, Bowen Wu, Baoxun Wang, Zhuoran Wang, and Xiaolong Wang. 2018 · 2080
Closest in time.