2017

Relevance of Unsupervised Metrics in Task-Oriented Dialogue for Evaluating Natural Language Generation

Sharma, Shikhar, Asri, Layla El, Schulz, Hannes et al.

Understand

Automated metrics such as BLEU are widely used in the machine translation literature.

  • They have also been used recently in the dialogue community for evaluating dialogue response generation.
  • However, previous work in dialogue response generation has shown that these metrics do not correlate strongly with human judgment in the non task-oriented dialogue setting.
  • Task-oriented dialogue responses are expressed on narrower domains and exhibit lower diversity.

Reading the bibliography…