2023

Catwalk: A Unified Language Model Evaluation Framework for Many Datasets

Groeneveld, Dirk, Awadalla, Anas, Beltagy, Iz et al.

Understand

The success of large language models has shifted the evaluation paradigms in natural language processing (NLP).

  • The community's interest has drifted towards comparing NLP models across many tasks, domains, and datasets, often at an extreme scale.
  • This imposes new engineering challenges: efforts in constructing datasets and models have been fragmented, and their formats and interfaces are incompatible.
  • As a result, it often takes extensive (re)implementation efforts to make fair and controlled comparisons at scale.

Reading the bibliography…