2019

Judge the Judges: A Large-Scale Evaluation Study of Neural Language Models for Online Review Generation

Garbacea, Cristina, Carton, Samuel, Yan, Shiyan et al.

Understand

We conduct a large-scale, systematic study to evaluate the existing evaluation methods for natural language generation in the context of generating online product reviews.

  • We compare human-based evaluators with a variety of automated evaluation procedures, including discriminative evaluators that measure how well machine-generated text can be distinguished from human-written text, as well as word overlap metrics that assess how similar the generated text compares to human-written references.
  • We determine to what extent these different evaluators agree on the ranking of a dozen of state-of-the-art generators for online product reviews.
  • We find that human evaluators do not correlate well with discriminative evaluators, leaving a bigger question of whether adversarial accuracy is the correct objective for natural language generation.

Reading the bibliography…