Fetching the paper…
Reading the bibliography…
Human evaluation for natural language generation (NLG) often suffers from inconsistent user ratings.
The measurement of observer agreement for categorical data
J. Richard Landis and Gary G. Koch. 1977 · 1977
Earlier work this paper cites.
Magnitude estimation of linguistic acceptability
Ellen Gurman Bard, Dan Robertson, and Antonella Sorace. 1996 · 1996
Earlier work this paper cites.
Long short-term memory
Sepp Hochreiter and Jürgen Schmidhuber. 1997 · 1997
Earlier work this paper cites.
Statistical significance tests for machine translation evaluation
Philipp Koehn. 2004 · 2004
Earlier work this paper cites.
Trueskill TM : a Bayesian skill rating system
Ralf Herbrich, Tom Minka, and Thore Graepel. 2006 · 2006
Earlier work this paper cites.
(Meta-) evaluation of machine translation
Chris Callison-Burch, Cameron Fordyce, Philipp Koehn, Christof Monz, and Josh Schroeder. 2007 · 2007
Earlier work this paper cites.
Individual and domain adaptation in sentence planning for dialogue
Marilyn Walker, Amanda Stent, François Mairesse, and Rashmi Prasad. 2007 · 2007
Earlier work this paper cites.
Sequence-to-sequence generation for spoken dialogue via deep syntax trees and strings
Ondřej Dušek and Filip Jurčíček. 2016 · 2008
Earlier work this paper cites.
Comparing rating scales and preference judgements in language evaluation
Anja Belz and Eric Kow. 2010 · 2010
Earlier work this paper cites.
Statistical machine translation
Philipp Koehn. 2010 · 2010
Earlier work this paper cites.
Blogging birds: Generating narratives about reintroduced species to promote public engagement
Advaith Siddharthan, Matthew Green, Kees van Deemter, Chris Mellish, and René van der Wal. 2012 · 2012
Earlier work this paper cites.
Offline sentence processing measures for testing readability with users
Advaith Siddharthan and Napoleon Katsos. 2012 · 2012
Earlier work this paper cites.
Continuous Measurement Scales in Human Evaluation of Machine Translation
Yvette Graham, Timothy Baldwin, Alistair Moffat, and Justin Zobel. 2013 · 2013
Cited alongside, same era.
Cluster-based prediction of user ratings for stylistic surface realisation
Nina Dethlefs, Heriberto Cuayáhuitl, Helen Hastie, Verena Rieser, and Oliver Lemon. 2014 · 2014
Cited alongside, same era.
A snapshot of NLG evaluation practices 2005–2014
Dimitra Gkatzia and Saad Mahamood. 2015 · 2014
Cited alongside, same era.
Natural language generation as incremental planning under uncertainty: Adaptive information presentation for statistical dialogue systems
Verena Rieser, Oliver Lemon, and Simon Keizer. 2014 · 2014
Cited alongside, same era.
Efficient elicitation of annotations for human evaluation of machine translation
Keisuke Sakaguchi, Matt Post, and Benjamin Van Durme. 2014 · 2014
Cited alongside, same era.
Natural language generation in dialogue using lexicalized and delexicalized data
Shikhar Sharma, Jing He, Kaheer Suleman, Hannes Schulz, and Philip Bachman. 2016 · 2016
Later among the works it cites.
Findings of the 2017 conference on machine translation (WMT17)
Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Shujian Huang, Matthias Huck, Philipp Koehn, Qun Liu, Varvara Logacheva, et al. 2017 · 2017
Later among the works it cites.
Referenceless Quality Estimation for Natural Language Generation
Ondřej Dušek, Jekaterina Novikova, and Verena Rieser. 2017 · 2017
Later among the works it cites.
Albert Gatt and Emiel Krahmer. 2017 · 2017
Later among the works it cites.
Towards an automatic turing test: Learning to evaluate dialogue responses
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Dzmitry Bahdanau, Kyunghyun Cho, and Yoshua Bengio. 2015 · 2015
Cited alongside, same era.
Evaluating human pairwise preference judgements
Mark Dras. 2015 · 2015
Cited alongside, same era.
Semantically conditioned LSTM-based natural language generation for spoken dialogue systems
Tsung-Hsien Wen, Milica Gašić, Nikola Mrkšić, Pei-Hao Su, David Vandyke, and Steve Young. 2015b · 2015
Cited alongside, same era.
Findings of the 2016 conference on machine translation (WMT16)
Ondřej Bojar, Rajen Chatterjee, Christian Federmann, Yvette Graham, Barry Haddow, Matthias Huck, Antonio Jimeno Yepes, Philipp Koehn, Varvara Logacheva, Christof Monz, Matteo Negri, Aurelie Neveol, Mariana Neves, Martin Popel, Matt Post, Raphael Rubino, Carolina Scarton, Lucia Specia, Marco Turchi, Karin Verspoor, and Marcos Zampieri. 2016 · 2016
Cited alongside, same era.
Automatic corpus extension for data-driven natural language generation
Elena Manishina, Bassam Jabaian, Stéphane Huet, and Fabrice Lefevre. 2016 · 2016
Cited alongside, same era.
Crowd-sourcing NLG data: Pictures elicit better data
Jekaterina Novikova, Oliver Lemon, and Verena Rieser. 2016 · 2016
Cited alongside, same era.
The E2E dataset: New challenges for end-to-end generation
Jekaterina Novikova, Ondřej Dušek, and Verena Rieser. 2017b
Cited in the paper.
Ryan Lowe, Michael Noseworthy, Iulian Vlad Serban, Nicolas Angelard-Gontier, Yoshua Bengio, and Joelle Pineau. 2017 · 2017
Later among the works it cites.
Why we need new evaluation metrics for NLG
Jekaterina Novikova, Ondřej Dušek, Amanda Cercas Curry, and Verena Rieser. 2017a · 2017
Later among the works it cites.
Challenges in data-to-document generation
Sam Wiseman, Stuart M. Shieber, and Alexander M. Rush. 2017 · 2017
Later among the works it cites.
Sheffield at E2E: structured prediction approaches to end-to-end language generation
Mingjie Chen, Gerasimos Lampouras, and Andreas Vlachos. 2018 · 2018
Closest in time.
A Deep Ensemble Model with Slot Alignment for Sequence-to-Sequence Natural Language Generation
Juraj Juraska, Panagiotis Karagiannis, Kevin K. Bowden, and Marilyn A. Walker. 2018 · 2018
Closest in time.
Discrete vs. continuous rating scales for language evaluation in NLP
Anja Belz and Eric Kow. 2011 · 2040
Closest in time.
Natural language generation enhances human decision-making with uncertain information
Dimitra Gkatzia, Oliver Lemon, and Verena Rieser. 2016 · 2043
Closest in time.