2023

What's In My Big Data?

Elazar, Yanai, Bhagia, Akshita, Magnusson, Ian et al.

Understand

Large text corpora are the backbone of language models.

  • However, we have a limited understanding of the content of these corpora, including general statistics, quality, social factors, and inclusion of evaluation data (contamination).
  • In this work, we propose What's In My Big Data? (WIMBD), a platform and a set of sixteen analyses that allow us to reveal and compare the contents of large text corpora.
  • WIMBD builds on two basic capabilities -- count and search -- at scale, which allows us to analyze more than 35 terabytes on a standard compute node.

Reading the bibliography…