2020

Establishing Baselines for Text Classification in Low-Resource Languages

Cruz, Jan Christian Blaise, Cheng, Charibeth

Understand

While transformer-based finetuning techniques have proven effective in tasks that involve low-resource, low-data environments, a lack of properly established baselines and benchmark datasets make it hard to compare different approaches that are aimed at tackling the low-resource setting.

  • In this work, we provide three contributions.
  • First, we introduce two previously unreleased datasets as benchmark datasets for text classification and low-resource multilabel text classification for the low-resource language Filipino.
  • Second, we pretrain better BERT and DistilBERT models for use within the Filipino setting.

Reading the bibliography…