2020

The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Gao, Leo, Biderman, Stella, Black, Sid et al.

Understand

Recent work has demonstrated that increased training dataset diversity improves general cross-domain knowledge and downstream generalization capability for large-scale language models.

  • With this in mind, we present \textit{the Pile}: an 825 GiB English text corpus targeted at training large-scale language models.
  • The Pile is constructed from 22 diverse high-quality subsets -- both existing and newly constructed -- many of which derive from academic or professional sources.
  • Our evaluation of the untuned performance of GPT-2 and GPT-3 on the Pile shows that these models struggle on many of its components, such as academic writing.

Reading the bibliography…