2020

AI4Bharat-IndicNLP Corpus: Monolingual Corpora and Word Embeddings for Indic Languages

Kunchukuttan, Anoop, Kakwani, Divyanshu, Golla, Satish et al.

Understand

We present the IndicNLP corpus, a large-scale, general-domain corpus containing 2.7 billion words for 10 Indian languages from two language families.

  • We share pre-trained word embeddings trained on these corpora.
  • We create news article category classification datasets for 9 languages to evaluate the embeddings.
  • We show that the IndicNLP embeddings significantly outperform publicly available pre-trained embedding on multiple evaluation tasks.

Reading the bibliography…