Fetching the paper…
Reading the bibliography…
We study the problem of multilingual masked language modeling, i.e.
Cross-lingual language model pretraining
Guillaume Lample and Alexis Conneau. 2019 · 1901
Earlier work this paper cites.
Wikimatrix: Mining 135m parallel sentences in 1620 language pairs from wikipedia
Holger Schwenk, Vishrav Chaudhary, Shuo Sun, Hongyu Gong, and Francisco Guzman. 2019 · 1907
Earlier work this paper cites.
On the cross-lingual transferability of monolingual representations
Mikel Artetxe, Sebastian Ruder, and Dani Yogatama. 2019 · 1910
Earlier work this paper cites.
Are girls neko or shōjo? cross-lingual alignment of non-isomorphic embeddings with iterative normalization
Mozhi Zhang, Keyulu Xu, Ken-ichi Kawarabayashi, Stefanie Jegelka, and Jordan Boyd-Graber. 2019 · 1911
Earlier work this paper cites.
“cloze procedure”: A new tool for measuring readability
Wilson L Taylor. 1953 · 1953
Earlier work this paper cites.
Cross-lingual name tagging and linking for 282 languages
Xiaoman Pan, Boliang Zhang, Jonathan May, Joel Nothman, Kevin Knight, and Heng Ji. 2017 · 1958
Earlier work this paper cites.
Adversarial training for unsupervised bilingual lexicon induction
Meng Zhang, Yang Liu, Huanbo Luan, and Maosong Sun. 2017 · 1970
Earlier work this paper cites.
Content and cluster analysis: assessing representational similarity in neural systems
Aarre Laakso and Garrison Cottrell. 2000 · 2000
Earlier work this paper cites.
Introduction to the CoNLL-2003 shared task: Language-independent named entity recognition
Erik F. Tjong Kim Sang and Fien De Meulder. 2003 · 2003
Earlier work this paper cites.
The third international Chinese language processing bakeoff: Word segmentation and named entity recognition
Gina-Anne Levow. 2006 · 2006
Earlier work this paper cites.
Moses: Open source toolkit for statistical machine translation
Philipp Koehn, Hieu Hoang, Alexandra Birch, Chris Callison-Burch, Marcello Federico, Nicola Bertoldi, Brooke Cowan, Wade Shen, Christine Moran, Richard Zens, Chris Dyer, Ondřej Bojar, Alexandra Constantin, and Evan Herbst. 2007 · 2007
Earlier work this paper cites.
A simple, fast, and effective reparameterization of IBM model 2
Chris Dyer, Victor Chahuneau, and Noah A. Smith. 2013 · 2013
Earlier work this paper cites.
Exploiting similarities among languages for machine translation
Tomas Mikolov, Quoc V Le, and Ilya Sutskever. 2013 · 2013
Earlier work this paper cites.
PanLex: Building a resource for panlingual lexical translation
David Kamholz, Jonathan Pool, and Susan Colowick. 2014 · 2014
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Earlier work this paper cites.
Deep biaffine attention for neural dependency parsing
Timothy Dozat and Christopher D Manning. 2016 · 2016
Earlier work this paper cites.
Gaussian error linear units (gelus)
Dan Hendrycks and Kevin Gimpel. 2016 · 2016
Earlier work this paper cites.
Convergent learning: Do different neural networks learn the same representations?
Yixuan Li, Jason Yosinski, Jeff Clune, Hod Lipson, and John E Hopcroft. 2016 · 2016
Cited alongside, same era.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016 · 2016
Cited alongside, same era.
Learning bilingual word embeddings with (almost) no bilingual data
Mikel Artetxe, Gorka Labaka, and Eneko Agirre. 2017 · 2017
Cited alongside, same era.
Enriching word vectors with subword information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017 · 2017
Cited alongside, same era.
Word translation without parallel data
Alexis Conneau, Guillaume Lample, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou. 2017 · 2017
Cited alongside, same era.
On the limitations of unsupervised bilingual dictionary induction
Anders Søgaard, Sebastian Ruder, and Ivan Vulić. 2018 · 2018
Later among the works it cites.
Towards understanding learning representations: To what extent do different neural networks learn the same representation
Liwei Wang, Lunjia Hu, Jiayuan Gu, Zhiqiang Hu, Yue Wu, Kun He, and John Hopcroft. 2018 · 2018
Later among the works it cites.
Massively multilingual neural machine translation
Roee Aharoni, Melvin Johnson, and Orhan Firat. 2019 · 2019
Closest in time.
Massively multilingual sentence embeddings for zero-shot cross-lingual transfer and beyond
Mikel Artetxe and Holger Schwenk. 2019 · 2019
Closest in time.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Ofir Press and Lior Wolf. 2017 · 2017
Cited alongside, same era.
Svcca: Singular vector canonical correlation analysis for deep learning dynamics and interpretability
Maithra Raghu, Justin Gilmer, Jason Yosinski, and Jascha Sohl-Dickstein. 2017 · 2017
Cited alongside, same era.
Offline bilingual word vectors, orthogonal transformations and the inverted softmax
Samuel L Smith, David HP Turban, Steven Hamblin, and Nils Y Hammerla. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
XNLI: Evaluating cross-lingual sentence representations
Alexis Conneau, Ruty Rinott, Guillaume Lample, Adina Williams, Samuel Bowman, Holger Schwenk, and Veselin Stoyanov. 2018 · 2018
Cited alongside, same era.
Universal language model fine-tuning for text classification
Jeremy Howard and Sebastian Ruder. 2018 · 2018
Cited alongside, same era.
Phrase-based & neural unsupervised machine translation
Guillaume Lample, Myle Ott, Alexis Conneau, Ludovic Denoyer, and Marc’Aurelio Ranzato. 2018 · 2018
Cited alongside, same era.
Haoyang Huang, Yaobo Liang, Nan Duan, Ming Gong, Linjun Shou, Daxin Jiang, and Ming Zhou. 2019 · 2019
Closest in time.
Similarity of neural network representations revisited
Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey Hinton. 2019 · 2019
Closest in time.
Investigating multilingual NMT representations at scale
Sneha Kudugunta, Ankur Bapna, Isaac Caswell, and Orhan Firat. 2019 · 2019
Closest in time.
Polyglot contextual representations improve crosslingual transfer
Phoebe Mulcaire, Jungo Kasai, and Noah A. Smith. 2019 · 2019
Closest in time.
Analyzing the limitations of cross-lingual word embedding mappings
Aitor Ormazabal, Mikel Artetxe, Gorka Labaka, Aitor Soroa, and Eneko Agirre. 2019 · 2019
Closest in time.
Bilingual lexicon induction with semi-supervision in non-isometric embedding spaces
Barun Patra, Joel Ruben Antony Moniz, Sarthak Garg, Matthew R. Gormley, and Graham Neubig. 2019 · 2019
Closest in time.
How multilingual is multilingual BERT?
Telmo Pires, Eva Schlinger, and Dan Garrette. 2019 · 2019
Closest in time.
Cross-lingual alignment of contextual word embeddings, with applications to zero-shot dependency parsing
Tal Schuster, Ori Ram, Regina Barzilay, and Amir Globerson. 2019 · 2019
Closest in time.
Cross-lingual BERT transformation for zero-shot dependency parsing
Yuxuan Wang, Wanxiang Che, Jiang Guo, Yijia Liu, and Ting Liu. 2019 · 2019
Closest in time.
Beto, bentz, becas: The surprising cross-lingual effectiveness of BERT
Shijie Wu and Mark Dredze. 2019 · 2019
Closest in time.
Cross-lingual ability of multilingual bert: An empirical study
K Karthikeyan, Zihan Wang, Stephen Mayhew, and Dan Roth. 2020 · 2020
Closest in time.