Fetching the paper…
Reading the bibliography…
This paper presents an effective approach for parallel corpus mining using bilingual sentence embeddings.
Mining the web for bilingual text
Philip Resnik. 1999 · 1999
Earlier work this paper cites.
Parallel web text mining for cross-language ir
Jiang Chen and Jian-Yun Nie. 2000 · 2000
Earlier work this paper cites.
Discriminative training and maximum entropy models for statistical machine translation
Franz Josef Och and Hermann Ney. 2002 · 2002
Earlier work this paper cites.
Bleu: a method for automatic evaluation of machine translation
Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002 · 2002
Earlier work this paper cites.
Mining english/chinese parallel documents from the world wide web
Christopher C Yang and Kar Wing Li. 2002 · 2002
Earlier work this paper cites.
The web as a parallel corpus
Philip Resnik and Noah A Smith. 2003 · 2003
Earlier work this paper cites.
Reliable measures for aligning japanese-english news articles and sentences
Masao Utiyama and Hitoshi Isahara. 2003 · 2003
Earlier work this paper cites.
Improving machine translation performance by exploiting non-parallel corpora
Dragos Stefan Munteanu and Daniel Marcu. 2005 · 2005
Earlier work this paper cites.
Extracting parallel sub-sentential fragments from non-parallel corpora
Dragos Stefan Munteanu and Daniel Marcu. 2006 · 2006
Earlier work this paper cites.
A dom tree alignment model for mining parallel data from the web
Lei Shi, Cheng Niu, Ming Zhou, and Jianfeng Gao. 2006 · 2006
Earlier work this paper cites.
Moses: Open source toolkit for statistical machine translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondřej Bojar, Alexandra Constantin, and Evan Herbst. 2007 · 2007
Earlier work this paper cites.
Mining a comparable text corpus for a vietnamese-french statistical machine translation system
Thi-Ngoc-Diep Do, Viet-Bac Le, Brigitte Bigi, Laurent Besacier, and Eric Castelli. 2009 · 2009
Cited alongside, same era.
Large scale parallel document mining for machine translation
Jakob Uszkoreit, Jay M. Ponte, Ashok C. Popat, and Moshe Dubiner. 2010 · 2010
Cited alongside, same era.
Building a web-based parallel corpus and filtering out machine-translated text
Alexandra Antonova and Alexey Misyurev. 2011 · 2011
Cited alongside, same era.
Semeval-2012 task 6: A pilot on semantic textual similarity
Eneko Agirre, Daniel Cer, Mona Diab, and Aitor Gonzalez-Agirre. 2012 · 2012
Cited alongside, same era.
Japanese and korean voice search
M. Schuster and K. Nakajima. 2012 · 2012
Cited alongside, same era.
Findings of the 2013 Workshop on Statistical Machine Translation
Ondřej Bojar, Christian Buck, Chris Callison-Burch, Christian Federmann, Barry Haddow, Philipp Koehn, Christof Monz, Matt Post, Radu Soricut, and Lucia Specia. 2013 · 2013
A deep neural network approach to parallel sentence extraction
Francis Grégoire and Philippe Langlais. 2017 · 2017
Later among the works it cites.
Efficient natural language response suggestion for smart reply
Matthew Henderson, Rami Al-Rfou, Brian Strope, Yun-Hsuan Sung, László Lukács, Ruiqi Guo, Sanjiv Kumar, Balint Miklos, and Ray Kurzweil. 2017 · 2017
Later among the works it cites.
Billion-scale similarity search with gpus
Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2017 · 2017
Later among the works it cites.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Later among the works it cites.
Zipporah: a fast and scalable data cleaning system for noisy web-crawled parallel corpora
Hainan Xu and Philipp Koehn. 2017 · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Nearest neighbor search in google correlate
Dan Vanderkam, Rob Schonberger, Henry Rowley, and Sanjiv Kumar. 2013 · 2013
Cited alongside, same era.
Findings of the 2014 workshop on statistical machine translation
Ondřej Bojar, Christian Buck, Christian Federmann, Barry Haddow, Philipp Koehn, Johannes Leveling, Christof Monz, Pavel Pecina, Matt Post, Herve Saint-Amand, et al. 2014 · 2014
Cited alongside, same era.
Deep unordered composition rivals syntactic methods for text classification
Mohit Iyyer, Varun Manjunatha, Jordan Boyd-Graber, and Hal Daumé III. 2015 · 2015
Cited alongside, same era.
Abcnn: Attention-based convolutional neural network for modeling sentence pairs
Wenpeng Yin, Hinrich Schütze, Bing Xiang, and Bowen Zhou. 2015 · 2015
Cited alongside, same era.
The united nations parallel corpus v1. 0
Michal Ziemski, Marcin Junczys-Dowmunt, and Bruno Pouliquen. 2016 · 2016
Cited alongside, same era.
Later among the works it cites.
H2@bucc18: Parallel sentence extraction from comparable corpora using multilingual sentence embeddings
Houda Bouamor and Hassan Sajjad. 2018 · 2018
Closest in time.
Daniel Cer, Yinfei Yang, Sheng-yi Kong, Nan Hua, Nicole Limtiaco, Rhomni St. John, Noah Constant, Mario Guajardo-Cespedes, Steve Yuan, Chris Tar, Yun-Hsuan Sung, Brian Strope, and Ray Kurzweil. 2018 · 2018
Closest in time.
Achieving human parity on automatic chinese to english news translation
Hany Hassan, Anthony Aue, Chang Chen, Vishal Chowdhary, Jonathan Clark, Christian Federmann, Xuedong Huang, Marcin Junczys-Dowmunt, William Lewis, Mu Li, et al. 2018 · 2018
Closest in time.
Filtering and mining parallel data in a joint multilingual space
Holger Schwenk. 2018 · 2018
Closest in time.
Learning semantic textual similarity from conversations
Yinfei Yang, Steve Yuan, Daniel Cer, Sheng-yi Kong, Noah Constant, Petr Pilar, Heming Ge, Yun-Hsuan Sung, Brian Strope, and Ray Kurzweil. 2018 · 2018
Closest in time.