2020

OCNLI: Original Chinese Natural Language Inference

Hu, Hai, Richardson, Kyle, Xu, Liang et al.

Understand

Despite the tremendous recent progress on natural language inference (NLI), driven largely by large-scale investment in new datasets (e.g., SNLI, MNLI) and advances in modeling, most progress has been limited to English due to a lack of reliable datasets for most of the world's languages.

  • In this paper, we present the first large-scale NLI dataset (consisting of ~56,000 annotated sentence pairs) for Chinese called the Original Chinese Natural Language Inference dataset (OCNLI).
  • Unlike recent attempts at extending NLI to other languages, our dataset does not rely on any automatic translation or non-expert annotation.
  • Instead, we elicit annotations from native speakers specializing in linguistics.

Reading the bibliography…