2024

A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains

Jacovi, Alon, Bitton, Yonatan, Bohnet, Bernd et al.

Understand

Prompting language models to provide step-by-step answers (e.g., "Chain-of-Thought") is the prominent approach for complex reasoning tasks, where more accurate reasoning chains typically improve downstream task performance.

  • Recent literature discusses automatic methods to verify reasoning to evaluate and improve their correctness.
  • However, no fine-grained step-level datasets are available to enable thorough evaluation of such verification methods, hindering progress in this direction.
  • We introduce REVEAL: Reasoning Verification Evaluation, a dataset to benchmark automatic verifiers of complex Chain-of-Thought reasoning in open-domain question-answering settings.

Reading the bibliography…