2022

M2D2: A Massively Multi-domain Language Modeling Dataset

Reid, Machel, Zhong, Victor, Gururangan, Suchin et al.

Understand

We present M2D2, a fine-grained, massively multi-domain corpus for studying domain adaptation in language models (LMs).

  • M2D2 consists of 8.5B tokens and spans 145 domains extracted from Wikipedia and Semantic Scholar.
  • Using ontologies derived from Wikipedia and ArXiv categories, we organize the domains in each data source into 22 groups.
  • This two-level hierarchy enables the study of relationships between domains and their effects on in- and out-of-domain performance after adaptation.

Reading the bibliography…