Fetching the paper…
Reading the bibliography…
This paper presents an automatic method to evaluate the naturalness of natural language generation in dialogue systems.
Machine translation evaluation with bert regressor
Hiroki Shimanaka, Tomoyuki Kajiwara, and Mamoru Komachi. 2019 · 1907
Earlier work this paper cites.
Least squares support vector machine classifiers
Johan AK Suykens and Joos Vandewalle. 1999 · 1999
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Rouge: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Satanjeev Banerjee and Alon Lavie. 2005 · 2005
Earlier work this paper cites.
Evaluating evaluation methods for generation in the presence of variation
Amanda Stent, Matthew Marge, and Mohit Singhai. 2005 · 2005
Earlier work this paper cites.
Comparing automatic and human evaluation of nlg systems
Anja Belz and Ehud Reiter. 2006 · 2006
Earlier work this paper cites.
Evaluation of text generation: A survey
Asli Celikyilmaz, Elizabeth Clark, and Jianfeng Gao. 2020 · 2006
Earlier work this paper cites.
A survey on transfer learning
Sinno Jialin Pan and Qiang Yang. 2009 · 2009
Earlier work this paper cites.
Empirical methods in natural language generation: Data-oriented methods and empirical evaluation , volume 5790
Emiel Krahmer and Mariët Theune. 2010 · 2010
Earlier work this paper cites.
Phrase-based statistical language generation using graphical models and active learning
François Mairesse, Milica Gašić, Filip Jurčíček, Simon Keizer, Blaise Thomson, Kai Yu, and Steve Young. 2010 · 2010
Cited alongside, same era.
Understanding bag-of-words model: a statistical framework
Yin Zhang, Rong Jin, and Zhi-Hua Zhou. 2010 · 2010
Cited alongside, same era.
Comparison of values of pearson’s and spearman’s correlation coefficients on the same sets of data
Jan Hauke and Tomasz Kossowski. 2011 · 2011
Cited alongside, same era.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Cited alongside, same era.
Semantically conditioned lstm-based natural language generation for spoken dialogue systems
TH Wen, M Gašić, N Mrkšić, PH Su, D Vandyke, and S Young. 2015 · 2015
Cited alongside, same era.
The price of debiasing automatic metrics in natural language evalaution
Arun Tejasvi Chaganty, Stephen Mussmann, and Percy Liang. 2018 · 2018
Later among the works it cites.
Survey of the state of the art in natural language generation: Core tasks, applications and evaluation
Albert Gatt and Emiel Krahmer. 2018 · 2018
Later among the works it cites.
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. 2018 · 2018
Later among the works it cites.
A structured review of the validity of bleu
Ehud Reiter. 2018 · 2018
Later among the works it cites.
Bert: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Unifying human and statistical evaluation for natural language generation
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Crowd-sourcing nlg data: Pictures elicit better data
Jekaterina Novikova, Oliver Lemon, and Verena Rieser. 2016 · 2016
Cited alongside, same era.
Referenceless quality estimation for natural language generation
Ondrej Dusek, Jekaterina Novikova, and Verena Rieser. 2017 · 2017
Cited alongside, same era.
Why we need new evaluation metrics for nlg
Jekaterina Novikova, Ondřej Dušek, Amanda Cercas Curry, and Verena Rieser. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
Tatsunori Hashimoto, Hugh Zhang, and Percy Liang. 2019 · 2019
Later among the works it cites.
Cost-sensitive bert for generalisable sentence classification on imbalanced data
Harish Tayyar Madabushi, Elena Kochkina, and Michael Castelle. 2019 · 2019
Later among the works it cites.
Bertscore: Evaluating text generation with bert
Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. 2019 · 2019
Later among the works it cites.
Bleurt: Learning robust metrics for text generation
Thibault Sellam, Dipanjan Das, and Ankur Parikh. 2020 · 2020
Later among the works it cites.