Fetching the paper…
Reading the bibliography…
We present an approach based on multilingual sentence embeddings to automatically extract parallel sentences from the content of Wikipedia articles in 85 languages, including several dialects or low-resource languages.
Yinfei Yang, Gustavo Hernández Ábrego, Steve Yuan, Mandy Guo, Qinlan Shen, Daniel Cer, Yun-Hsuan Sung, Brian Strope, and Ray Kurzweil. 2019 · 1902
Earlier work this paper cites.
Mining the Web for Bilingual Text
Philip Resnik. 1999 · 1999
Earlier work this paper cites.
The Web as a Parallel Corpus
Philip Resnik and Noah A. Smith. 2003 · 2003
Earlier work this paper cites.
Reliable Measures for Aligning Japanese-English News Articles and Sentences
Masao Utiyama and Hitoshi Isahara. 2003 · 2003
Earlier work this paper cites.
Europarl: A parallel corpus for statistical machine translation
Philipp Koehn. 2005 · 2005
Earlier work this paper cites.
Improving Machine Translation Performance by Exploiting Non-Parallel Corpora
Dragos Stefan Munteanu and Daniel Marcu. 2005 · 2005
Earlier work this paper cites.
Finding similar sentences across multiple languages in Wikipedia
Sisay Fissaha Adafre and Maarten de Rijke. 2006 · 2006
Earlier work this paper cites.
On the Use of Comparable Corpora to Improve SMT performance
Sadaf Abdul-Rauf and Holger Schwenk. 2009 · 2009
Earlier work this paper cites.
Building bilingual parallel corpora based on Wikipedia
Mehdi Zadeh Mohammadi and Nasser GhasemAghaee. 2010 · 2010
Earlier work this paper cites.
Wikipedia as multilingual source of comparable corpora
Pablo Gamallo Otero and Isaac González López. 2010 · 2010
Earlier work this paper cites.
Extracting parallel sentences from comparable corpora using document level alignment
Jason R. Smith, Chris Quirk, and Kristina Toutanova. 2010 · 2010
Earlier work this paper cites.
Product quantization for nearest neighbor search
H. Jégou, M. Douze, and C. Schmid. 2011 · 2011
Earlier work this paper cites.
Measuring comparability of multilingual corpora extracted from Wikipedia
P Otero, I López, S Cilenis, and Santiago de Compostela. 2011 · 2011
Earlier work this paper cites.
Identifying parallel documents from a large bilingual collection of texts: Application to parallel article extraction in Wikipedia
Alexandre Patry and Philippe Langlais. 2011 · 2011
Earlier work this paper cites.
Parallel data, tools and interfaces in OPUS
J. Tiedemann. 2012 · 2012
Cited alongside, same era.
Wikipedia as an smt training corpus
Dan Tufis, Radu Ion, Ștefan Daniel, Dumitrescu, and Dan Ștefănescu. 2013 · 2013
Cited alongside, same era.
Findings of the wmt 2016 bilingual document alignment shared task
Christian Buck and Philipp Koehn. 2016 · 2016
Cited alongside, same era.
Bag of tricks for efficient text classification
Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. 2016 · 2016
Cited alongside, same era.
Opensubtitles2016: Extracting large parallel corpora from movie and tv subtitles
P. Lison and J. Tiedemann. 2016 · 2016
Cited alongside, same era.
Cross-lingual wikification using multilingual embeddings
Chen-Tse Tsai and Dan Roth. 2016 · 2016
Extracting Parallel Sentences from Comparable Corpora with STACC Variants
Andoni Azpeitia, Thierry Etchegoyhen, and Eva Martínez Garcia. 2018 · 2018
Later among the works it cites.
H2@BUCC18: Parallel Sentence Extraction from Comparable Corpora Using Multilingual Sentence Embeddings
Houda Bouamor and Hassan Sajjad. 2018 · 2018
Later among the works it cites.
Set-Theoretic Alignment for Comparable Corpora
Thierry Etchegoyhen and Andoni Azpeitia. 2016 · 2018
Later among the works it cites.
Effective Parallel Corpus Mining using Bilingual Sentence Embeddings
Mandy Guo, Qinlan Shen, Yinfei Yang, Heming Ge, Daniel Cer, Gustavo Hernandez Abrego, Keith Stevens, Noah Constant, Yun-Hsuan Sung, Brian Strope, and Ray Kurzweil. 2018 · 2018
Later among the works it cites.
Achieving Human Parity on Automatic Chinese to English News Translation
Hany Hassan, Anthony Aue, Chang Chen, Vishal Chowdhary, Jonathan Clark, Christian Federmann, Xuedong Huang, Marcin Junczys-Dowmunt, William Lewis, Mu Li, Shujie Liu, Tie-Yan Liu, Renqian Luo, Arul Menezes, Tao Qin, Frank Seide, Xu Tan, Fei Tian, Lijun Wu, Shuangzhi Wu, Yingce Xia, Dongdong Zhang, Zhirui Zhang, and Ming Zhou. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
The United Nations Parallel Corpus v1.0
Michał Ziemski, Marcin Junczys-Dowmunt, and Bruno Pouliquen. 2016 · 2016
Cited alongside, same era.
Weighted Set-Theoretic Alignment of Comparable Sentences
Andoni Azpeitia, Thierry Etchegoyhen, and Eva Martínez Garcia. 2017 · 2017
Cited alongside, same era.
An Empirical Analysis of NMT-Derived Interlingual Embeddings and their Use in Parallel Sentence Identification
Cristina España-Bonet, Ádám Csaba Varga, Alberto Barrón-Cedeño, and Josef van Genabith. 2017 · 2017
Cited alongside, same era.
Multiwiki: Interlingual text passage alignment in Wikipedia
Simon Gottschalk and Elena Demidova. 2017 · 2017
Cited alongside, same era.
BUCC 2017 Shared Task: a First Attempt Toward a Deep Learning Framework for Identifying Parallel Sentences in Comparable Corpora
Francis Grégoire and Philippe Langlais. 2017 · 2017
Cited alongside, same era.
Billion-scale similarity search with GPUs
Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2017 · 2017
Cited alongside, same era.
Later among the works it cites.
Findings of the wmt 2018 shared task on parallel corpus filtering
Philipp Koehn, Huda Khayrallah, Kenneth Heafield, and Mikel L. Forcada. 2018 · 2018
Later among the works it cites.
Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing
Taku Kudo and John Richardson. 2018 · 2018
Later among the works it cites.
A call for clarity in reporting bleu scores
Matt Post. 2018 · 2018
Later among the works it cites.
When and why are pre-trained word embeddings useful for neural machine translation?
Ye Qi, Devendra Sachan, Matthieu Felix, Sarguna Padmanabhan, and Graham Neubig. 2018 · 2018
Later among the works it cites.
Filtering and mining parallel data in a joint multilingual space
Holger Schwenk. 2018 · 2018
Later among the works it cites.
Low-resource corpus filtering using multilingual sentence embeddings
Vishrav Chaudhary, Yuqing Tang, Francisco Guzmán, Holger Schwenk, and Philipp Koehn. 2019 · 2019
Closest in time.
Findings of the wmt 2019 shared task on parallel corpus filtering for low-resource conditions
Philipp Koehn, Francisco Guzmán, Vishrav Chaudhary, and Juan M. Pino. 2019 · 2019
Closest in time.
fairseq: A fast, extensible toolkit for sequence modeling
Myle Ott, Sergey Edunov, Alexei Baevski, Angela Fan, Sam Gross, Nathan Ng, David Grangier, and Michael Auli. 2019 · 2019
Closest in time.