Fetching the paper…
Reading the bibliography…
We learn a joint multilingual sentence embedding and use the distance between sentences in different languages to filter noisy parallel data and to mine for parallel data in large news collections.
The web as a parallel corpus
Philip Resnik and Noah A. Smith. 2003 · 2003
Earlier work this paper cites.
Reliable measures for aligning Japanese-English news articles and sentences
Masao Utiyama and Hitoshi Isahara. 2003 · 2003
Earlier work this paper cites.
Improving machine translation performance by exploiting non-parallel corpora
Dragos Stefan Munteanu and Daniel Marcu. 2005 · 2005
Earlier work this paper cites.
On the use of comparable corpora to improve SMT performance
Sadaf Abdul Rauf and Holger Schwenk. 2009 · 2009
Earlier work this paper cites.
Domain adaptation via pseudo in-domain data selection
Amittai Axelrod, Xiaodong He, and Jianfeng Gao. 2011 · 2011
Earlier work this paper cites.
Multilingual deep learning
Sarath Chandar, Mitesh M. Khapra, Balaraman Ravindran, Vikas Raykar, and Amrita Saha. 2013 · 2013
Earlier work this paper cites.
Multilingual models for compositional distributed semantics
Karl Moritz Hermann and Phil Blunsom. 2014 · 2014
Earlier work this paper cites.
Effective approaches to attention-based neural machine translation
Minh-Thang Luong, Hieu Pham, and Christophe D. Manning. 2015 · 2015
Earlier work this paper cites.
Learning distributed representations for multilingual text sequences
Hieu Pham, Minh-Thang Luong, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
A character-level decoder without explicit segmentation for neural machine translation
Junyoung Chunga, Kyunghyun Choa, and Yoshua Bengio. 2016 · 2016
Earlier work this paper cites.
Set-theoretic alignment of comparable corpora
Thierry Etchegoyhen and Andoni Azpeitia. 2016 · 2016
Cited alongside, same era.
Google’s multilingual neural machine translation system: Enabling zero-shot translation
Melvin Johnson et al. 2016 · 2016
Cited alongside, same era.
Bilingual word embeddings from parallel and non-parallel corpora for cross-language classification
Aditua Mogadala and Achim Rettinger. 2016 · 2016
Cited alongside, same era.
Google’s neural machine translation system: Bridging the gap between human and machine translation
Yonghui Wu et al. 2016 · 2016
Cited alongside, same era.
Cross-lingual sentiment classification with bilingual document representation learning
Xinjie Zhou, Xiaojun Wan, and Jianguo Xiao. 2016 · 2016
Cited alongside, same era.
Billion-scale similarity search with gpus
Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2017 · 2017
Later among the works it cites.
Data selection with cluster-based language difference models and cynical selection
Lucía Santamaría and Amittai Axelrod. 2017 · 2017
Later among the works it cites.
Learning joint multilingual sentence representations with neural machine translation
Holger Schwenk and Matthijs Douze. 2017 · 2017
Later among the works it cites.
Dynamic data selection for neural machine translation
Marlies van der Wees, Arianna Bisazza, and Christof Monz. 2017 · 2017
Later among the works it cites.
Extracting parallel sentences from comparable corpora with STACC variants
Andoni Azpeitia, Thierry Etchegoyhen, and Eva Martínez Garcia. 2018 · 2018
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Attention is all you need
Ashish Vaswani et al. 2017 · 2017
Cited alongside, same era.
Synthetic and artificial noise both break neural machine translation
Yonatan Belinkov and Yonatan Bisk. 2017 · 2017
Cited alongside, same era.
An empirical analysis of nmt-derived interlingual embeddings and their use in parallel sentence identification
Cristina España-Bonet, Ádám Csaba Varga, Alberto Barrón-Cedeño, and Josef van Genabith. 2017 · 2017
Cited alongside, same era.
Convolutional Sequence to Sequence Learning
Jonas Gehring, Michael Auli, David Grangier, Denis Yarats, and Yann N Dauphin. 2017 · 2017
Cited alongside, same era.
Bucc 2017 shared task: a first attempt toward a deep learning framework for identifying parallel sentences in comparable corpora
Francis Grégoire and Philippe Langlais. 2017 · 2017
Cited alongside, same era.
Improving neural machine translation models with monolingual data
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016a
Cited in the paper.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016b
Cited in the paper.
H2@bucc18: Parallel sentence extraction from comparable coprora using multlingual sentence embeddings
Houda Bouamor and Hassan Sajjad. 2018 · 2018
Closest in time.
Achieving human parity on automatic chinese to english translation
Hany Hassan, Anthony Aue, Chang Chen, Vishal Chowdhary, Jonathan Clark, Christian Federmann, Xuedong Huang, Marcin Junczys-Dowmunt, William Lewis, Mu Li, Shujie Liu, Tie-Yan Liu, Renqian Luo, Arul Menezes, Tao Qin, Frank Seide, Xu Tan, Fei Tian, Lijun Wu, Shuangzhi Wu, Yingce Xia, Dongdong Zhang, Zhirui Zhang, and Ming Zhou. 2018 · 2018
Closest in time.
Um- p p aligner: Neural network-based parallel sentence identification model
Chongman Leong, Derek F. Wong, and Lidia S. Chao. 2018 · 2018
Closest in time.
Overview of the third bucc shared task: Spottign parallel sentences in comparable corpora
Pierre Zweigenbaum, Serge Sharoff, and Reinhard Rapp. 2018 · 2018
Closest in time.