Fetching the paper…
Reading the bibliography…
We introduce an architecture to learn joint multilingual sentence representations for 93 languages, belonging to more than 30 different families and written in 28 different scripts.
Cross-lingual language model pretraining
Guillaume Lample and Alexis Conneau. 2019 · 1901
Earlier work this paper cites.
RCV1: A new benchmark collection for text categorization research
David D. Lewis, Yiming Yang, Tony G. Rose, and Fan Li. 2004 · 2004
Earlier work this paper cites.
Inducing crosslingual distributed representations of words
Alexandre Klementiev, Ivan Titov, and Binod Bhattarai. 2012 · 2012
Earlier work this paper cites.
Parallel data, tools and interfaces in OPUS
Jörg Tiedemann. 2012 · 2012
Earlier work this paper cites.
Distributed representations of sentences and documents
Quoc Le and Tomas Mikolov. 2014 · 2014
Earlier work this paper cites.
GloVe: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Earlier work this paper cites.
A large annotated corpus for learning natural language inference
Samuel R. Bowman, Gabor Angeli, Christopher Potts, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Semi-supervised sequence learning
Andrew M. Dai and Quoc V. Le. 2015 · 2015
Earlier work this paper cites.
BilBOWA: Fast bilingual distributed representations without word alignments
Stephan Gouws, Yoshua Bengio, and Greg Corrado. 2015 · 2015
Earlier work this paper cites.
Skip-thought vectors
Ryan Kiros, Yukun Zhu, Ruslan R. Salakhutdinov, Richard Zemel, Raquel Urtasun, Antonio Torralba, and Sanja Fidler. 2015 · 2015
Earlier work this paper cites.
Bilingual word representations with monolingual quality in mind
Thang Luong, Hieu Pham, and Christopher D. Manning. 2015 · 2015
Earlier work this paper cites.
Massively multilingual word embeddings
Waleed Ammar, George Mulcaire, Yulia Tsvetkov, Guillaume Lample, Chris Dyer, and Noah A. Smith. 2016 · 2016
Earlier work this paper cites.
Improving neural machine translation models with monolingual data
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Earlier work this paper cites.
Weighted set-theoretic alignment of comparable sentences
Andoni Azpeitia, Thierry Etchegoyhen, and Eva Martínez Garcia. 2017 · 2017
Earlier work this paper cites.
Supervised learning of universal sentence representations from natural language inference data
Alexis Conneau, Douwe Kiela, Holger Schwenk, Loïc Barrault, and Antoine Bordes. 2017 · 2017
Earlier work this paper cites.
An Empirical Analysis of NMT-Derived Interlingual Embeddings and their Use in Parallel Sentence Identification
Cristina España-Bonet, Ádám Csaba Varga, Alberto Barrón-Cedeño, and Josef van Genabith. 2017 · 2017
Earlier work this paper cites.
BUCC 2017 shared task: a first attempt toward a deep learning framework for identifying parallel sentences in comparable corpora
Francis Grégoire and Philippe Langlais. 2017 · 2017
Cited alongside, same era.
Google’s multilingual neural machine translation system: Enabling zero-shot translation
Melvin Johnson, Mike Schuster, Quoc V. Le, Maxim Krikun, Yonghui Wu, Zhifeng Chen, Nikhil Thorat, Fernanda Viégas, Martin Wattenberg, Greg Corrado, Macduff Hughes, and Jeffrey Dean. 2017 · 2017
Cited alongside, same era.
Learning language representations for typology prediction
Chaitanya Malaviya, Graham Neubig, and Patrick Littell. 2017 · 2017
Cited alongside, same era.
A survey of cross-lingual embedding models
Sebastian Ruder, Ivan Vulic, and Anders Søgaard. 2017 · 2017
Cited alongside, same era.
Learning joint multilingual sentence representations with neural machine translation
Holger Schwenk and Matthijs Douze. 2017 · 2017
Cited alongside, same era.
Embedding learning through multilingual concept induction
Philipp Dufter, Mengjie Zhao, Martin Schmitt, Alexander Fraser, and Hinrich Schütze. 2018 · 2018
Closest in time.
Understanding back-translation at scale
Sergey Edunov, Myle Ott, Michael Auli, and David Grangier. 2018 · 2018
Closest in time.
Effective parallel corpus mining using bilingual sentence embeddings
Mandy Guo, Qinlan Shen, Yinfei Yang, Heming Ge, Daniel Cer, Gustavo Hernandez Abrego, Keith Stevens, Noah Constant, Yun-Hsuan Sung, Brian Strope, and Ray Kurzweil. 2018 · 2018
Closest in time.
Achieving human parity on automatic Chinese to English news translation
Hany Hassan, Anthony Aue, Chang Chen, Vishal Chowdhary, Jonathan Clark, Christian Federmann, Xuedong Huang, Marcin Junczys-Dowmunt, William Lewis, Mu Li, Shujie Liu, Tie-Yan Liu, Renqian Luo, Arul Menezes, Tao Qin, Frank Seide, Xu Tan, Fei Tian, Lijun Wu, Shuangzhi Wu, Yingce Xia, Dongdong Zhang, Zhirui Zhang, and Ming Zhou. 2018 · 2018
Closest in time.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder. 2018 · 2018
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
zNLP: Identifying parallel sentences in Chinese-English comparable corpora
Zheng Zhang and Pierre Zweigenbaum. 2017 · 2017
Cited alongside, same era.
Overview of the second BUCC shared task: Spotting parallel sentences in comparable corpora
Pierre Zweigenbaum, Serge Sharoff, and Reinhard Rapp. 2017 · 2017
Cited alongside, same era.
Unsupervised statistical machine translation
Mikel Artetxe, Gorka Labaka, and Eneko Agirre. 2018c · 2018
Cited alongside, same era.
Margin-based parallel corpus mining with multilingual sentence embeddings
Mikel Artetxe and Holger Schwenk. 2018 · 2018
Cited alongside, same era.
Extracting Parallel Sentences from Comparable Corpora with STACC Variants
Andoni Azpeitia, Thierry Etchegoyhen, and Eva Martínez Garcia. 2018 · 2018
Cited alongside, same era.
H2@BUCC18: Parallel Sentence Extraction from Comparable Corpora Using Multilingual Sentence Embeddings
Houda Bouamor and Hassan Sajjad. 2018 · 2018
Cited alongside, same era.
Closest in time.
Phrase-based & neural unsupervised machine translation
Guillaume Lample, Myle Ott, Alexis Conneau, Ludovic Denoyer, and Marc’Aurelio Ranzato. 2018 · 2018
Closest in time.
Rapid adaptation of neural machine translation to new languages
Graham Neubig and Junjie Hu. 2018 · 2018
Closest in time.
Deep contextualized word representations
Matthew Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer. 2018 · 2018
Closest in time.
Filtering and mining parallel data in a joint multilingual space
Holger Schwenk. 2018 · 2018
Closest in time.
A corpus for multilingual document classification in eight languages
Holger Schwenk and Xian Li. 2018 · 2018
Closest in time.
Learning general purpose distributed sentence representations via large scale multi-task learning
Sandeep Subramanian, Adam Trischler, Yoshua Bengio, and Christopher J. Pal. 2018 · 2018
Closest in time.
GLUE: A multi-task benchmark and analysis platform for natural language understanding
Alex Wang, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. 2018 · 2018
Closest in time.
A broad-coverage challenge corpus for sentence understanding through inference
Adina Williams, Nikita Nangia, and Samuel Bowman. 2018 · 2018
Closest in time.
Multilingual seq2seq training with similarity loss for cross-lingual document classification
Katherine Yu, Haoran Li, and Barlas Oguz. 2018 · 2018
Closest in time.
Overview of the Third BUCC Shared Task: Spotting Parallel Sentences in Comparable Corpora
Pierre Zweigenbaum, Serge Sharoff, and Reinhard Rapp. 2018 · 2018
Closest in time.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Closest in time.