2022

TRUE: Re-evaluating Factual Consistency Evaluation

Honovich, Or, Aharoni, Roee, Herzig, Jonathan et al.

Understand

Grounded text generation systems often generate text that contains factual inconsistencies, hindering their real-world applicability.

  • Automatic factual consistency evaluation may help alleviate this limitation by accelerating evaluation cycles, filtering inconsistent outputs and augmenting training data.
  • While attracting increasing attention, such evaluation metrics are usually developed and evaluated in silo for a single task or dataset, slowing their adoption.
  • Moreover, previous meta-evaluation protocols focused on system-level correlations with human annotations, which leave the example-level accuracy of such metrics unclear.

Reading the bibliography…