2021

BanglaBERT: Language Model Pretraining and Benchmarks for Low-Resource Language Understanding Evaluation in Bangla

Bhattacharjee, Abhik, Hasan, Tahmid, Ahmad, Wasi Uddin et al.

Understand

In this work, we introduce BanglaBERT, a BERT-based Natural Language Understanding (NLU) model pretrained in Bangla, a widely spoken yet low-resource language in the NLP literature.

  • To pretrain BanglaBERT, we collect 27.5 GB of Bangla pretraining data (dubbed `Bangla2B+') by crawling 110 popular Bangla sites.
  • We introduce two downstream task datasets on natural language inference and question answering and benchmark on four diverse NLU tasks covering text classification, sequence labeling, and span prediction.
  • In the process, we bring them under the first-ever Bangla Language Understanding Benchmark (BLUB).

Reading the bibliography…