Fetching the paper…
Reading the bibliography…
Open-domain human-computer conversation has been attracting increasing attention over the past few years.
PARADISE: A framework for evaluating spoken dialogue agents
Marilyn A Walker, Diane J Litman, Candace A Kamm, and Alicia Abella. 1997 · 1997
Earlier work this paper cites.
Quantitative and qualitative evaluation of darpa communicator spoken dialogue systems
Marilyn A Walker, Rebecca Passonneau, and Julie E Boland. 2001 · 2001
Earlier work this paper cites.
Automatic evaluation of machine translation quality using n-gram co-occurrence statistics
George Doddington. 2002 · 2002
Earlier work this paper cites.
BLEU: A method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Establishing and maintaining long-term human-computer relationships
Timothy W Bickmore and Rosalind W Picard. 2005 · 2005
Earlier work this paper cites.
Quantitative evaluation of user simulation techniques for spoken dialogue systems
Jost Schatzmann, Kallirroi Georgila, and Steve Young. 2005 · 2005
Earlier work this paper cites.
Re-evaluation the role of BLEU in machine translation research
Chris Callison-Burch, Miles Osborne, and Philipp Koehn. 2006 · 2006
Earlier work this paper cites.
Regression for sentence-level MT evaluation with pseudo references
Joshua Albrecht and Rebecca Hwa. 2007 · 2007
Earlier work this paper cites.
Heterogeneous automatic MT evaluation through non-parametric metric combinations
Jesús Giménez and Lluís Márquez. 2008 · 2008
Earlier work this paper cites.
Estimating the sentence-level quality of machine translation systems
Lucia Specia, Marco Turchi, Nicola Cancedda, Marc Dymetman, and Nello Cristianini. 2009 · 2009
Earlier work this paper cites.
Machine translation evaluation versus quality estimation
Lucia Specia, Dhwaj Raj, and Marco Turchi. 2010 · 2010
Cited alongside, same era.
Evaluate with confidence estimation: Machine ranking of translation outputs using grammatical features
Eleftherios Avramidis, Maja Popović, David Vilar, and Aljoscha Burchardt. 2011 · 2011
Cited alongside, same era.
Evaluation without references: IBM1 scores as evaluation metrics
Maja Popović, David Vilar, Eleftherios Avramidis, and Aljoscha Burchardt. 2011 · 2011
Cited alongside, same era.
Data-driven response generation in social media
Alan Ritter, Colin Cherry, and William B Dolan. 2011 · 2011
Cited alongside, same era.
Dialog system using real-time crowdsourcing and twitter large-scale corpus
Fumihiro Bessho, Tatsuya Harada, and Yasuo Kuniyoshi. 2012 · 2012
Cited alongside, same era.
Automatically assessing machine summary content without a gold standard
Neural responding machine for short-text conversation
Lifeng Shang, Zhengdong Lu, and Hang Li. 2015 · 2015
Later among the works it cites.
A neural network approach to context-sensitive generation of conversational responses
Alessandro Sordoni, Michel Galley, Michael Auli, Chris Brockett, Yangfeng Ji, Margaret Mitchell, Jian-Yun Nie, Jianfeng Gao, and Bill Dolan. 2015 · 2015
Later among the works it cites.
Results of the WMT15 metrics shared task
Miloš Stanojević, Amirand Koehn Philipp Kamran, and Ondřej Bojar. 2015 · 2015
Later among the works it cites.
Results of the WMT16 metrics shared task
Ondřej Bojar, Yvette Graham, Amir Kamran, and Miloš Stanojević. 2016 · 2016
Later among the works it cites.
A diversity-promoting objective function for neural conversation models
Jiwei Li, Michel Galley, Chris Brockett, Jianfeng Gao, and Bill Dolan. 2015 · 2016
Later among the works it cites.
How NOT to evaluate your dialogue system: An empirical study of unsupervised evaluation metrics for dialogue response generation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Annie Louis and Ani Nenkova. 2013 · 2013
Cited alongside, same era.
Bootstrapping dialog systems with word embeddings
Gabriel Forgues, Joelle Pineau, Jean-Marie Larchevêque, and Réal Tremblay. 2014 · 2014
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik Kingma and Jimmy Ba. 2014 · 2014
Cited alongside, same era.
Sequence to sequence learning with neural networks
Ilya Sutskever, Oriol Vinyals, and Quoc V Le. 2014 · 2014
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Cited alongside, same era.
deltaBLEU: a discriminative metric for generation tasks with intrinsically diverse targets
Michel Galley, Chris Brockett, Alessandro Sordoni, Yangfeng Ji, Michael Auli, Chris Quirk, Margaret Mitchell, Jianfeng Gao, and Bill Dolan. 2015 · 2015
Cited alongside, same era.
Learning to rank short text pairs with convolutional deep neural networks
Aliaksei Severyn and Alessandro Moschitti. 2015 · 2015
Cited alongside, same era.
Chia-Wei Liu, Ryan Lowe, Iulian Serban, Mike Noseworthy, Laurent Charlin, and Joelle Pineau. 2016 · 2016
Later among the works it cites.
Sequence to backward and forward sequences: A content-introducing approach to generative short-text conversation
Lili Mou, Yiping Song, Rui Yan, Ge Li, Lu Zhang, and Zhi Jin. 2016b · 2016
Later among the works it cites.
A hierarchical latent variable encoder-decoder model for generating dialogues
Iulian Vlad Serban, Alessandro Sordoni, Ryan Lowe, Laurent Charlin, Joelle Pineau, Aaron Courville, and Yoshua Bengio. 2016 · 2016
Later among the works it cites.
Two are better than one: An ensemble of retrieval-and generation-based dialog systems
Yiping Song, Rui Yan, Xiang Li, Dongyan Zhao, and Ming Zhang. 2016 · 2016
Later among the works it cites.
Learning natural language inference with LSTM
Shuohang Wang and Jing Jiang. 2016 · 2016
Later among the works it cites.
Learning to respond with deep neural networks for retrieval-based human-computer conversation system
Rui Yan, Yiping Song, and Hua Wu. 2016 · 2016
Later among the works it cites.
Towards an automatic Turing test: Learning to evaluate dialogue responses
Ryan Lowe, Michael Noseworthy, Iulian V. Serban, Nicolas Angelard-Gontier, Yoshua Bengio, and Joelle Pineau. 2017 · 2017
Closest in time.