2021

Evaluating Attribution in Dialogue Systems: The BEGIN Benchmark

Dziri, Nouha, Rashkin, Hannah, Linzen, Tal et al.

Understand

Knowledge-grounded dialogue systems powered by large language models often generate responses that, while fluent, are not attributable to a relevant source of information.

  • Progress towards models that do not exhibit this issue requires evaluation metrics that can quantify its prevalence.
  • To this end, we introduce the Benchmark for Evaluation of Grounded INteraction (BEGIN), comprised of 12k dialogue turns generated by neural dialogue systems trained on three knowledge-grounded dialogue corpora.
  • We collect human annotations assessing the extent to which the models' responses can be attributed to the given background information.

Reading the bibliography…