Fetching the paper…
Reading the bibliography…
An important aspect of developing dialogue systems is how to evaluate and compare the performance of different systems.
The second conversational intelligence challenge (convai2)
Emily Dinan, Varvara Logacheva, Valentin Malykh, Alexander Miller, Kurt Shuster, Jack Urbanek, Douwe Kiela, Arthur Szlam, Iulian Serban, Ryan Lowe, et al. 2019 · 1902
Earlier work this paper cites.
Better automatic evaluation of open-domain dialogue systems with contextualized embeddings
Sarik Ghazarian, Johnny Tian-Zheng Wei, A. Galstyan, and Nanyun Peng. 2019 · 1904
Earlier work this paper cites.
Roberta: A robustly optimized bert pretraining approach
Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1907
Earlier work this paper cites.
Stackgan++: Realistic image synthesis with stacked generative adversarial networks
Han Zhang, Tao Xu, Hongsheng Li, Shaoting Zhang, Xiaogang Wang, Xiaolei Huang, and Dimitris N Metaxas. 2018a · 1962
Earlier work this paper cites.
The fréchet distance between multivariate normal distributions
DC Dowson and BV Landau. 1982 · 1982
Earlier work this paper cites.
The measurement of textual coherence with latent semantic analysis
Peter W Foltz, Walter Kintsch, and Thomas K Landauer. 1998 · 1998
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, S. Roukos, T. Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Vector-based models of semantic composition
Jeff Mitchell and Mirella Lapata. 2008 · 2008
Earlier work this paper cites.
Power comparisons of shapiro-wilk, kolmogorov-smirnov, lilliefors and anderson-darling tests
Nornadiah Mohd Razali, Yap Bee Wah, et al. 2011 · 2011
Earlier work this paper cites.
An optimal assessment of natural language student input using word-to-word similarity metrics
Vasile Rus and Mihai Lintean. 2012 · 2012
Earlier work this paper cites.
Meteor universal: Language specific translation evaluation for any target language
Michael Denkowski and Alon Lavie. 2014 · 2014
Earlier work this paper cites.
Bootstrapping dialog systems with word embeddings
Gabriel Forgues, Joelle Pineau, Jean-Marie Larchevêque, and Réal Tremblay. 2014 · 2014
Cited alongside, same era.
Oriol Vinyals and Quoc Le. 2015 · 2015
Cited alongside, same era.
Towards universal paraphrastic sentence embeddings
John Wieting, Mohit Bansal, Kevin Gimpel, and Karen Livescu. 2015 · 2015
Cited alongside, same era.
Gans trained by a two time-scale update rule converge to a local nash equilibrium
Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. 2017 · 2017
Cited alongside, same era.
Progressive growing of gans for improved quality, stability, and variation
Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. 2017 · 2017
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Semantic image synthesis with spatially-adaptive normalization
Taesung Park, Ming-Yu Liu, Ting-Chun Wang, and Jun-Yan Zhu. 2019 · 2019
Later among the works it cites.
Xlnet: Generalized autoregressive pretraining for language understanding
Zhilin Yang, Zihang Dai, Yiming Yang, Jaime Carbonell, Russ R Salakhutdinov, and Quoc V Le. 2019 · 2019
Later among the works it cites.
Electra: Pre-training text encoders as discriminators rather than generators
Kevin Clark, Minh-Thang Luong, Quoc V Le, and Christopher D Manning. 2020 · 2020
Later among the works it cites.
Grade: Automatic graph-enhanced coherence metric for evaluating open-domain dialogue systems
Lishan Huang, Zheng Ye, Jinghui Qin, Liang Lin, and Xiaodan Liang. 2020 · 2020
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Towards an automatic turing test: Learning to evaluate dialogue responses
Ryan Lowe, Michael Noseworthy, I. Serban, Nicolas Angelard-Gontier, Yoshua Bengio, and Joelle Pineau. 2017 · 2017
Cited alongside, same era.
Parlai: A dialog research software platform
Alexander H Miller, Will Feng, Adam Fisch, Jiasen Lu, Dhruv Batra, Antoine Bordes, Devi Parikh, and Jason Weston. 2017 · 2017
Cited alongside, same era.
Learning discourse-level diversity for neural dialog models using conditional variational autoencoders
Tiancheng Zhao, Ran Zhao, and Maxine Eskenazi. 2017 · 2017
Cited alongside, same era.
Mechanism-aware neural machine for dialogue response generation
Ganbin Zhou, Ping Luo, Rongyu Cao, Fen Lin, Bo Chen, and Qing He. 2017 · 2017
Cited alongside, same era.
Towards less generic responses in neural conversation models: A statistical re-weighting method
Yahui Liu, Wei Bi, Jun Gao, Xiaojiang Liu, Jian Yao, and Shuming Shi. 2018 · 2018
Cited alongside, same era.
Towards empathetic open-domain conversation models: A new benchmark and dataset
Hannah Rashkin, Eric Michael Smith, Margaret Li, and Y-Lan Boureau. 2018 · 2018
Cited alongside, same era.
Assessing generative models via precision and recall
Mehdi SM Sajjadi, Olivier Bachem, Mario Lucic, Olivier Bousquet, and Sylvain Gelly. 2018 · 2018
Cited alongside, same era.
Usr: An unsupervised and reference free evaluation metric for dialog generation
Shikib Mehri and M. Eskénazi. 2020 · 2020
Later among the works it cites.
Bleurt: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Later among the works it cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2020 · 2020
Later among the works it cites.
Designing precise and robust dialogue response evaluators
Tianyu Zhao, Divesh Lala, and Tatsuya Kawahara. 2020 · 2020
Later among the works it cites.
Enhancing the open-domain dialogue evaluation in latent space
Zhangming Chan, Lemao Liu, Juntao Li, Haisong Zhang, Dongyan Zhao, Shuming Shi, and Rui Yan. 2021 · 2021
Closest in time.
A training-free and reference-free summarization evaluation metric via centrality-weighted relevance and self-referenced redundancy
Wang Chen, Piji Li, and Irwin King. 2021 · 2021
Closest in time.
REAM ♯ \mbox{REAM}\sharp : An enhancement approach to reference-based evaluation metrics for open-domain dialog generation
Jun Gao, Wei Bi, Ruifeng Xu, and Shuming Shi. 2021 · 2021
Closest in time.