Fetching the paper…
Reading the bibliography…
Automatic evaluation of open-domain dialogue response generation is very challenging because there are many appropriate responses for a given context.
Measuring nominal scale agreement among many raters
Joseph L Fleiss. 1971 · 1971
Earlier work this paper cites.
Feature normalization and likelihood-based similarity measures for image retrieval
Selim Aksoy and Robert M Haralick. 2001 · 2001
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
The probabilistic relevance framework: Bm25 and beyond
Stephen Robertson, Hugo Zaragoza, et al. 2009 · 2009
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen. 2010 · 2010
Earlier work this paper cites.
What is left to be understood in atis?
Gokhan Tur, Dilek Hakkani-Tür, and Larry Heck. 2010 · 2010
Earlier work this paper cites.
Chameleons in imagined conversations: A new approach to understanding coordination of linguistic style in dialogs
Cristian Danescu-Niculescu-Mizil and Lillian Lee. 2011 · 2011
Earlier work this paper cites.
Metrics and evaluation of spoken dialogue systems
Helen Hastie. 2012 · 2012
Earlier work this paper cites.
A fast and simple algorithm for training neural probabilistic language models
Andriy Mnih and Yee Whye Teh. 2012 · 2012
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013 · 2013
Earlier work this paper cites.
A systematic comparison of smoothing techniques for sentence-level bleu
Boxing Chen and Colin Cherry. 2014 · 2014
Cited alongside, same era.
Glove: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher D. Manning. 2014 · 2014
Cited alongside, same era.
Learning fine-grained image similarity with deep ranking
J. Wang, Y. Song, T. Leung, C. Rosenberg, J. Wang, J. Philbin, B. Chen, and Y. Wu. 2014 · 2014
Cited alongside, same era.
Facenet: A unified embedding for face recognition and clustering
F. Schroff, D. Kalenichenko, and J. Philbin. 2015 · 2015
Cited alongside, same era.
A hierarchical recurrent encoder-decoder for generative context-aware query suggestion
Alessandro Sordoni, Yoshua Bengio, Hossein Vahabi, Christina Lioma, Jakob Grue Simonsen, and Jian-Yun Nie. 2015 · 2015
Cited alongside, same era.
How not to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation
Exemplar encoder-decoder for neural conversation generation
Gaurav Pandey, Danish Contractor, Vineet Kumar, and Sachindra Joshi. 2018 · 2018
Later among the works it cites.
A hierarchical latent structure for variational conversation modeling
Yookoon Park, Jaemin Cho, and Gunhee Kim. 2018 · 2018
Later among the works it cites.
Deep contextualized word representations
Matthew E. Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Later among the works it cites.
Ruber: An unsupervised method for automatic evaluation of open-domain dialog systems
Chongyang Tao, Lili Mou, Dongyan Zhao, and Rui Yan. 2018 · 2018
Later among the works it cites.
Variational hierarchical user-based conversation model
JinYeong Bak and Alice Oh. 2019 · 2019
Later among the works it cites.
Sentence mover’s similarity: Automatic evaluation for multi-sentence texts
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Chia-Wei Liu, Ryan Lowe, Iulian Serban, Mike Noseworthy, Laurent Charlin, and Joelle Pineau. 2016 · 2016
Cited alongside, same era.
Learning end-to-end goal-oriented dialog
Antoine Bordes, Y-Lan Boureau, and Jason Weston. 2017 · 2017
Cited alongside, same era.
Towards an automatic Turing test: Learning to evaluate dialogue responses
Ryan Lowe, Michael Noseworthy, Iulian Vlad Serban, Nicolas Angelard-Gontier, Yoshua Bengio, and Joelle Pineau. 2017 · 2017
Cited alongside, same era.
Why we need new evaluation metrics for NLG
Jekaterina Novikova, Ondřej Dušek, Amanda Cercas Curry, and Verena Rieser. 2017 · 2017
Cited alongside, same era.
A hierarchical latent variable encoder-decoder model for generating dialogues
Iulian Vlad Serban, Alessandro Sordoni, Ryan Lowe, Laurent Charlin, Joelle Pineau, Aaron C Courville, and Yoshua Bengio. 2017 · 2017
Cited alongside, same era.
An efficient framework for learning sentence representations
Lajanugen Logeswaran and Honglak Lee. 2018 · 2018
Cited alongside, same era.
Elizabeth Clark, Asli Celikyilmaz, and Noah A. Smith. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Investigating evaluation of open-domain dialogue systems with human generated multiple references
Prakhar Gupta, Shikib Mehri, Tiancheng Zhao, Amy Pavel, Maxine Eskenazi, and Jeffrey Bigham. 2019 · 2019
Later among the works it cites.
Sentence-BERT: Sentence embeddings using Siamese BERT-networks
Nils Reimers and Iryna Gurevych. 2019 · 2019
Later among the works it cites.
Re-evaluating adem: A deeper look at scoring dialogue responses
Ananya B. Sai, Mithun Das Gupta, Mitesh M. Khapra, and Mukundhan Srinivasan. 2019 · 2019
Later among the works it cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, and Yoav Artzi. 2020 · 2020
Closest in time.