Fetching the paper…
Reading the bibliography…
Large parallel corpora that are automatically obtained from the web, documents or elsewhere often exhibit many corrupted parts that are bound to negatively affect the quality of the systems and models that learn from these corpora.
S. Khadivi and H. Ney, Automatic Filtering of Bilingual Corpora for Statistical Machine Translation, in: Natural Language Processing and Information Systems, 10th International Conference on Applications of Natural Language to Information Systems
2005
Earlier work this paper cites.
P. Koehn, H. Hoang, A. Birch, C. Callison-Burch, M. Federico, N. Bertoldi, B. Cowan, W. Shen, C. Moran, R. Zens, C. Dyer, O. Bojar, A. Constantin and E. Herbst, Moses: open source toolkit for statistical machine translation, Association for Computational Linguistics, 2007, pp. 177–180. https://dl.acm.org/citation.cfm?id=1557821
2007
Earlier work this paper cites.
M. Lui and T. Baldwin, langid.py: An off-the-shelf language identification tool, in: Proceedings of the ACL 2012 System Demonstrations
2012
Earlier work this paper cites.
K. Wolk, Noisy-parallel and comparable corpora filtering methodology for the extraction of bi-lingual equivalent data at sentence level, Computer Science @BULLET Computer Science
2015
Cited alongside, same era.
2016
Cited alongside, same era.
H. Xu and P. Koehn, Zipporah: a Fast and Scalable Data Cleaning System for Noisy Web-Crawled Parallel Corpora, Emnlp
2017
Cited alongside, same era.
2017
Later among the works it cites.
M. Pinnis, M. Rikters and R. Krišlauks, Tilde’s Machine Translation Systems for WMT 2018, in: Proceedings of the Third Conference on Machine Translation (WMT 2018), Volume 2: Shared Task Papers
2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…