2018

The WiLI benchmark dataset for written language identification

Thoma, Martin

Understand

This paper describes the WiLI-2018 benchmark dataset for monolingual written natural language identification.

  • WiLI-2018 is a publicly available, free of charge dataset of short text extracts from Wikipedia.
  • It contains 1000 paragraphs of 235 languages, totaling in 23500 paragraphs.
  • WiLI is a classification dataset: Given an unknown paragraph written in one dominant language, it has to be decided which language it is.

Reading the bibliography…