2022

Towards a Unified Multi-Dimensional Evaluator for Text Generation

Zhong, Ming, Liu, Yang, Yin, Da et al.

Understand

Multi-dimensional evaluation is the dominant paradigm for human evaluation in Natural Language Generation (NLG), i.e., evaluating the generated text from multiple explainable dimensions, such as coherence and fluency.

  • However, automatic evaluation in NLG is still dominated by similarity-based metrics, and we lack a reliable framework for a more comprehensive evaluation of advanced models.
  • In this paper, we propose a unified multi-dimensional evaluator UniEval for NLG.
  • We re-frame NLG evaluation as a Boolean Question Answering (QA) task, and by guiding the model with different questions, we can use one evaluator to evaluate from multiple dimensions.

Reading the bibliography…