2022

Evaluate & Evaluation on the Hub: Better Best Practices for Data and Model Measurements

von Werra, Leandro, Tunstall, Lewis, Thakur, Abhishek et al.

Understand

Evaluation is a key part of machine learning (ML), yet there is a lack of support and tooling to enable its informed and systematic practice.

  • We introduce Evaluate and Evaluation on the Hub --a set of tools to facilitate the evaluation of models and datasets in ML.
  • Evaluate is a library to support best practices for measurements, metrics, and comparisons of data and models.
  • Its goal is to support reproducibility of evaluation, centralize and document the evaluation process, and broaden evaluation to cover more facets of model performance.

Reading the bibliography…