Fetching the paper…
Reading the bibliography…
Most of recent work in cross-lingual word embeddings is severely Anglocentric.
Learning unsupervised multilingual word embeddings with incremental multilingual hubs
Geert Heyman, Bregt Verreet, Ivan Vulić, and Marie-Francine Moens. 2019 · 1902
Earlier work this paper cites.
Wikimatrix: Mining 135m parallel sentences in 1620 language pairs from wikipedia
Holger Schwenk, Vishrav Chaudhary, Shuo Sun, Hongyu Gong, and Francisco Guzmán. 2019 · 1907
Earlier work this paper cites.
Adversarial training for unsupervised bilingual lexicon induction
Meng Zhang, Yang Liu, Huanbo Luan, and Maosong Sun. 2017 · 1970
Earlier work this paper cites.
Stanza: A python natural language processing toolkit for many human languages
Peng Qi, Yuhao Zhang, Yuhui Zhang, Jason Bolton, and Christopher D Manning. 2020 · 2003
Earlier work this paper cites.
Statistical significance tests for machine translation evaluation
Philipp Koehn. 2004 · 2004
Earlier work this paper cites.
Edinburgh system description for the 2005 iwslt speech translation evaluation
Philipp Koehn, Amittai Axelrod, Alexandra Birch Mayne, Chris Callison-Burch, Miles Osborne, and David Talbot. 2005 · 2005
Earlier work this paper cites.
Word alignment for languages with scarce resources using bilingual corpora of other language pairs
Haifeng Wang, Hua Wu, and Zhanyi Liu. 2006 · 2006
Earlier work this paper cites.
Statistical machine translation
Philipp Koehn. 2009 · 2009
Earlier work this paper cites.
Word representations: A simple and general method for semi-supervised learning
Joseph Turian, Lev-Arie Ratinov, and Yoshua Bengio. 2010 · 2010
Earlier work this paper cites.
What is morphology? , volume 8
Mark Aronoff and Kirsten Fudeman. 2011 · 2011
Earlier work this paper cites.
Inducing crosslingual distributed representations of words
Alexandre Klementiev, Ivan Titov, and Binod Bhattarai. 2012 · 2012
Earlier work this paper cites.
A simple, fast, and effective reparameterization of ibm model 2
Chris Dyer, Victor Chahuneau, and Noah A Smith. 2013 · 2013
Earlier work this paper cites.
Combining bilingual and comparable corpora for low resource machine translation
Ann Irvine and Chris Callison-Burch. 2013 · 2013
Earlier work this paper cites.
Exploiting similarities among languages for machine translation
Tomas Mikolov, Quoc V Le, and Ilya Sutskever. 2013 · 2013
Earlier work this paper cites.
Bilingual word embeddings for phrase-based machine translation
Will Y Zou, Richard Socher, Daniel Cer, and Christopher D Manning. 2013 · 2013
Cited alongside, same era.
How to make words with vectors: Phrase generation in distributional semantics
Georgiana Dinu and Marco Baroni. 2014 · 2014
Cited alongside, same era.
Improving zero-shot learning by mitigating the hubness problem
Georgiana Dinu, Angeliki Lazaridou, and Marco Baroni. 2015 · 2015
Cited alongside, same era.
Multi-task word alignment triangulation for low-resource languages
Tomer Levinboim and David Chiang. 2015 · 2015
Cited alongside, same era.
Normalized word embedding and orthogonal transform for bilingual word translation
Chao Xing, Dong Wang, Chao Liu, and Yiye Lin. 2015 · 2015
Cited alongside, same era.
Opensubtitles2015: Extracting large parallel corpora from movie and tv subtitles
Unsupervised multilingual word embeddings
Xilun Chen and Claire Cardie. 2018 · 2018
Later among the works it cites.
Word translation without parallel data
Alexis Conneau, Guillaume Lample, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou. 2018 · 2018
Later among the works it cites.
Learning word vectors for 157 languages
Edouard Grave, Piotr Bojanowski, Prakhar Gupta, Armand Joulin, and Tomas Mikolov. 2018 · 2018
Later among the works it cites.
Phrase-based & neural unsupervised machine translation
Guillaume Lample, Myle Ott, Alexis Conneau, Ludovic Denoyer, and Marc’Aurelio Ranzato. 2018 · 2018
Later among the works it cites.
When and why are pre-trained word embeddings useful for neural machine translation?
Ye Qi, Devendra Sachan, Matthieu Felix, Sarguna Padmanabhan, and Graham Neubig. 2018 · 2018
Later among the works it cites.
Jw300: A wide-coverage parallel corpus for low-resource languages
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Pierre Lison and Jörg Tiedemann. 2016 · 2016
Cited alongside, same era.
Universal dependencies v1: A multilingual treebank collection
Joakim Nivre, Marie-Catherine De Marneffe, Filip Ginter, Yoav Goldberg, Jan Hajic, Christopher D Manning, Ryan McDonald, Slav Petrov, Sampo Pyysalo, Natalia Silveira, et al. 2016 · 2016
Cited alongside, same era.
Ten pairs to tag–multilingual pos tagging via coarse mapping between embeddings
Yuan Zhang, David Gaddy, Regina Barzilay, and Tommi Jaakkola. 2016 · 2016
Cited alongside, same era.
Learning bilingual word embeddings with (almost) no bilingual data
Mikel Artetxe, Gorka Labaka, and Eneko Agirre. 2017 · 2017
Cited alongside, same era.
Uriel and lang2vec: Representing languages as typological, geographical, and phylogenetic vectors
Patrick Littell, David R Mortensen, Ke Lin, Katherine Kairis, Carlisle Turner, and Lori Levin. 2017 · 2017
Cited alongside, same era.
Learning language representations for typology prediction
Chaitanya Malaviya, Graham Neubig, and Patrick Littell. 2017 · 2017
Cited alongside, same era.
Offline bilingual word vectors, orthogonal transformations and the inverted softmax
Samuel L Smith, David HP Turban, Steven Hamblin, and Nils Y Hammerla. 2017 · 2017
Cited alongside, same era.
Željko Agić and Ivan Vulić. 2019 · 2019
Closest in time.
Unsupervised hyperalignment for multilingual word embeddings
Jean Alaux, Edouard Grave, Marco Cuturi, and Armand Joulin. 2019 · 2019
Closest in time.
Don’t forget the long tail! a comprehensive analysis of morphological generalization in bilingual lexicon induction
Paula Czarnowska, Sebastian Ruder, Edouard Grave, Ryan Cotterell, and Ann Copestake. 2019 · 2019
Closest in time.
How to (properly) evaluate cross-lingual word embeddings: On strong baselines, comparative analyses, and some misconceptions
Goran Glavaš, Robert Litschko, Sebastian Ruder, and Ivan Vulić. 2019 · 2019
Closest in time.
Lost in evaluation: Misleading benchmarks for bilingual dictionary induction
Yova Kementchedjhieva, Mareike Hartmann, and Anders Søgaard. 2019 · 2019
Closest in time.
Choosing transfer languages for cross-lingual learning
Yu-Hsiang Lin, Chian-Yu Chen, Jean Lee, Zirui Li, Yuyan Zhang, Mengzhou Xia, Shruti Rijhwani, Junxian He, Zhisong Zhang, Xuezhe Ma, Antonios Anastasopoulos, Patrick Littell, and Graham Neubig. 2019 · 2019
Closest in time.
Bilingual lexicon induction with semi-supervision in non-isometric embedding spaces
Barun Patra, Joel Ruben Antony Moniz, Sarthak Garg, Matthew R. Gormley, and Graham Neubig. 2019 · 2019
Closest in time.
Density matching for bilingual word embedding
Chunting Zhou, Xuezhe Ma, Di Wang, and Graham Neubig. 2019 · 2019
Closest in time.