2019

Bridging the domain gap in cross-lingual document classification

Lai, Guokun, Oguz, Barlas, Yang, Yiming et al.

Understand

The scarcity of labeled training data often prohibits the internationalization of NLP models to multiple languages.

  • Recent developments in cross-lingual understanding (XLU) has made progress in this area, trying to bridge the language barrier using language universal representations.
  • However, even if the language problem was resolved, models trained in one language would not transfer to another language perfectly due to the natural domain drift across languages and cultures.
  • We consider the setting of semi-supervised cross-lingual understanding, where labeled data is available in a source language (English), but only unlabeled data is available in the target language.

Reading the bibliography…