2022

ROSCOE: A Suite of Metrics for Scoring Step-by-Step Reasoning

Golovneva, Olga, Chen, Moya, Poff, Spencer et al.

Understand

Large language models show improved downstream task performance when prompted to generate step-by-step reasoning to justify their final answers.

  • These reasoning steps greatly improve model interpretability and verification, but objectively studying their correctness (independent of the final answer) is difficult without reliable methods for automatic evaluation.
  • We simply do not know how often the stated reasoning steps actually support the final end task predictions.
  • In this work, we present ROSCOE, a suite of interpretable, unsupervised automatic scores that improve and extend previous text generation evaluation metrics.

Reading the bibliography…