2021

Towards Standard Criteria for human evaluation of Chatbots: A Survey

Liang, Hongru, Li, Huaqing

Understand

Human evaluation is becoming a necessity to test the performance of Chatbots.

  • However, off-the-shelf settings suffer the severe reliability and replication issues partly because of the extremely high diversity of criteria.
  • It is high time to come up with standard criteria and exact definitions.
  • To this end, we conduct a through investigation of 105 papers involving human evaluation for Chatbots.

Reading the bibliography…