2021

What Will it Take to Fix Benchmarking in Natural Language Understanding?

Bowman, Samuel R., Dahl, George E.

Understand

Evaluation for many natural language understanding (NLU) tasks is broken: Unreliable and biased systems score so highly on standard benchmarks that there is little room for researchers who develop better systems to demonstrate their improvements.

  • The recent trend to abandon IID benchmarks in favor of adversarially-constructed, out-of-distribution test sets ensures that current models will perform poorly, but ultimately only obscures the abilities that we want our benchmarks to measure.
  • In this position paper, we lay out four criteria that we argue NLU benchmarks should meet.
  • We argue most current benchmarks fail at these criteria, and that adversarial data collection does not meaningfully address the causes of these failures.

Reading the bibliography…