2019

CCAligned: A Massive Collection of Cross-Lingual Web-Document Pairs

El-Kishky, Ahmed, Chaudhary, Vishrav, Guzman, Francisco et al.

Understand

Cross-lingual document alignment aims to identify pairs of documents in two distinct languages that are of comparable content or translations of each other.

  • In this paper, we exploit the signals embedded in URLs to label web documents at scale with an average precision of 94.5% across different language pairs.
  • We mine sixty-eight snapshots of the Common Crawl corpus and identify web document pairs that are translations of each other.
  • We release a new web dataset consisting of over 392 million URL pairs from Common Crawl covering documents in 8144 language pairs of which 137 pairs include English.

Reading the bibliography…