Fetching the paper…
Reading the bibliography…
The majority of NLG evaluation relies on automatic metrics, such as BLEU .
Evan James Williams. 1959 · 1959
Earlier work this paper cites.
The measurement of observer agreement for categorical data
J Richard Landis and Gary G Koch. 1977 · 1977
Earlier work this paper cites.
How to write plain English: A book for lawyers and consumers
Rudolf Franz Flesch. 1979 · 1979
Earlier work this paper cites.
Applying natural language generation to indicative summarization
Min-Yen Kan, Kathleen R. McKeown, and Judith L. Klavans. 2001 · 2001
Earlier work this paper cites.
Automatic evaluation of machine translation quality using n-gram co-occurrence statistics
George Doddington. 2002 · 2002
Earlier work this paper cites.
BLEU: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
ROUGE: A package for automatic evaluation of summaries
Chin-Yew Lin. 2004 · 2004
Earlier work this paper cites.
Evaluating evaluation methods for generation in the presence of variation
Amanda Stent, Matthew Marge, and Mohit Singhai. 2005 · 2005
Earlier work this paper cites.
Comparing automatic and human evaluation of NLG systems
Anja Belz and Ehud Reiter. 2006 · 2006
Earlier work this paper cites.
Re-evaluating the role of BLEU in machine translation research
Chris Callison-Burch, Miles Osborne, and Philipp Koehn. 2006 · 2006
Earlier work this paper cites.
A study of translation edit rate with targeted human annotation
Matthew Snover, Bonnie Dorr, Richard Schwartz, Linnea Micciulla, and John Makhoul. 2006 · 2006
Earlier work this paper cites.
METEOR: An automatic metric for MT evaluation with high levels of correlation with human judgments
Alon Lavie and Abhaya Agarwal. 2007 · 2007
Earlier work this paper cites.
Sequence-to-sequence generation for spoken dialogue via deep syntax trees and strings
Ondřej Dušek and Filip Jurčíček. 2016 · 2008
Earlier work this paper cites.
A smorgasbord of features for automatic MT evaluation
Jesús Giménez and Lluís Màrquez. 2008 · 2008
Earlier work this paper cites.
An investigation into the validity of some metrics for automatically evaluating natural language generation systems
Ehud Reiter and Anja Belz. 2009 · 2009
Earlier work this paper cites.
Further meta-evaluation of broad-coverage surface realization
Dominic Espinosa, Rajakrishnan Rajkumar, Michael White, and Shoshana Berleant. 2010 · 2010
Earlier work this paper cites.
Phrase-based statistical language generation using graphical models and active learning
François Mairesse, Milica Gašić, Filip Jurčíček, Simon Keizer, Blaise Thomson, Kai Yu, and Steve Young. 2010 · 2010
Cited alongside, same era.
Machine translation evaluation versus quality estimation
Lucia Specia, Dhwaj Raj, and Marco Turchi. 2010 · 2010
Cited alongside, same era.
UMBC_EBIQUITY-CORE: Semantic textual similarity systems
Lushan Han, Abhay Kashyap, Tim Finin, James Mayfield, and Jonathan Weese. 2013 · 2013
Cited alongside, same era.
Learning whom to trust with MACE
Dirk Hovy, Taylor Berg-Kirkpatrick, Ashish Vaswani, and Eduard H. Hovy. 2013 · 2013
Cited alongside, same era.
Recent Advances in Automatic Readability Assessment and Text Simplification
Thomas Francois and Delphine Bernhard, editors. 2014 · 2014
Cited alongside, same era.
A snapshot of NLG evaluation practices 2005–2014
Dimitra Gkatzia and Saad Mahamood. 2015 · 2014
What to talk about and how? Selective generation using LSTMs with coarse-to-fine alignment
Hongyuan Mei, Mohit Bansal, and Matthew R. Walter. 2016 · 2016
Later among the works it cites.
There’s no comparison: Reference-less evaluation metrics in grammatical error correction
Courtney Napoles, Keisuke Sakaguchi, and Joel Tetreault. 2016 · 2016
Later among the works it cites.
Crowd-sourcing NLG data: Pictures elicit better data
Jekaterina Novikova, Oliver Lemon, and Verena Rieser. 2016 · 2016
Later among the works it cites.
Natural language generation in dialogue using lexicalized and delexicalized data
Shikhar Sharma, Jing He, Kaheer Suleman, Hannes Schulz, and Philip Bachman. 2016 · 2016
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Natural language generation as incremental planning under uncertainty: Adaptive information presentation for statistical dialogue systems
Verena Rieser, Oliver Lemon, and Simon Keizer. 2014 · 2014
Cited alongside, same era.
Training a natural language generator from unaligned data
Ondřej Dušek and Filip Jurčíček. 2015 · 2015
Cited alongside, same era.
CIDEr: Consensus-based image description evaluation
Ramakrishna Vedantam, C. Lawrence Zitnick, and Devi Parikh. 2015 · 2015
Cited alongside, same era.
Semantically conditioned LSTM-based natural language generation for spoken dialogue systems
Tsung-Hsien Wen, Milica Gašić, Nikola Mrkšić, Pei-Hao Su, David Vandyke, and Steve Young. 2015 · 2015
Cited alongside, same era.
A context-aware natural language generator for dialogue systems
Ondřej Dušek and Filip Jurčíček. 2016 · 2016
Cited alongside, same era.
Why bother? Is evaluation of NLG in an end-to-end Spoken Dialogue System worth it?
Helen Hastie, Heriberto Cuayahuitl, Nina Dethlefs, Simon Keizer, and Xingkun Liu. 2016 · 2016
Cited alongside, same era.
Tsung-Hsien Wen, Milica Gašić, Nikola Mrkšić, Lina Maria Rojas-Barahona, Pei-hao Su, David Vandyke, and Steve J. Young. 2016 · 2016
Later among the works it cites.
Referenceless quality estimation for natural language generation
Ondrej Dušek, Jekaterina Novikova, and Verena Rieser. 2017 · 2017
Closest in time.
Adversarial evaluation of dialogue models
Anjuli Kannan and Oriol Vinyals. 2017 · 2017
Closest in time.
Re-evaluating automatic metrics for image captioning
Mert Kilickaya, Aykut Erdem, Nazli Ikizler-Cinbis, and Erkut Erdem. 2017 · 2017
Closest in time.
The E2E dataset: New challenges for end-to-end generation
Jekaterina Novikova, Ondrej Dušek, and Verena Rieser. 2017 · 2017
Closest in time.
Correlating human and automatic evaluation of a German surface realiser
Aoife Cahill. 2009 · 2025
Closest in time.
Discrete vs. continuous rating scales for language evaluation in NLP
Anja Belz and Eric Kow. 2011 · 2040
Closest in time.
Natural language generation enhances human decision-making with uncertain information
Dimitra Gkatzia, Oliver Lemon, and Verena Rieser. 2016 · 2043
Closest in time.
LEPOR: A robust evaluation metric for machine translation with augmented factors
Aaron L. F. Han, Derek F. Wong, and Lidia S. Chao. 2012 · 2044
Closest in time.
deltaBLEU: A discriminative metric for generation tasks with intrinsically diverse targets
Michel Galley, Chris Brockett, Alessandro Sordoni, Yangfeng Ji, Michael Auli, Chris Quirk, Margaret Mitchell, Jianfeng Gao, and Bill Dolan. 2015 · 2073
Closest in time.
Comparing automatic evaluation measures for image description
Desmond Elliott and Frank Keller. 2014 · 2074
Closest in time.