Fetching the paper…
Reading the bibliography…
Bilingual lexicons map words in one language to their translations in another, and are typically induced by learning linear projections to align monolingual word embedding spaces.
Adding interpretable attention to neural translation models improves word alignment
Thomas Zenkel, Joern Wuebker, and John DeNero. 2019 · 1901
Earlier work this paper cites.
Wikimatrix: Mining 135M parallel sentences in 1620 language pairs from Wikipedia
Holger Schwenk, Vishrav Chaudhary, Shuo Sun, Hongyu Gong, and Francisco Guzmán. 2019 · 1907
Earlier work this paper cites.
Unsupervised cross-lingual representation learning at scale
Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guillaume Wenzek, Francisco Guzmán, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. 2019 · 1911
Earlier work this paper cites.
The mathematics of statistical machine translation: Parameter estimation
Peter F. Brown, Stephen A. Della Pietra, Vincent J. Della Pietra, and Robert L. Mercer. 1993 · 1993
Earlier work this paper cites.
Mining the web for bilingual text
Philip Resnik. 1999 · 1999
Earlier work this paper cites.
Multilingual denoising pre-training for neural machine translation
Yinhan Liu, Jiatao Gu, Naman Goyal, Xian Li, Sergey Edunov, Marjan Ghazvininejad, Mike Lewis, and Luke Zettlemoyer. 2020 · 2001
Earlier work this paper cites.
A systematic comparison of various statistical alignment models
Franz Josef Och and Hermann Ney. 2003 · 2003
Earlier work this paper cites.
A DOM tree alignment model for mining parallel data from the web
Lei Shi, Cheng Niu, Ming Zhou, and Jianfeng Gao. 2006 · 2006
Earlier work this paper cites.
On the use of comparable corpora to improve SMT performance
Sadaf Abdul-Rauf and Holger Schwenk. 2009 · 2009
Earlier work this paper cites.
The wacky wide web: a collection of very large linguistically processed web-crawled corpora
Marco Baroni, Silvia Bernardini, Adriano Ferraresi, and Eros Zanchetta. 2009 · 2009
Earlier work this paper cites.
A simple, fast, and effective reparameterization of IBM model 2
Chris Dyer, Victor Chahuneau, and Noah A. Smith. 2013 · 2013
Earlier work this paper cites.
Exploiting similarities among languages for machine translation
Tomas Mikolov, Quoc V. Le, and Ilya Sutskever. 2013 · 2013
Earlier work this paper cites.
Improving zero-shot learning by mitigating the hubness problem
Georgiana Dinu, Angeliki Lazaridou, and Marco Baroni. 2015 · 2015
Earlier work this paper cites.
Adam: A method for stochastic optimization
Diederik P. Kingma and Jimmy Ba. 2015 · 2015
Cited alongside, same era.
Learning principled bilingual mappings of word embeddings while preserving monolingual invariance
Mikel Artetxe, Gorka Labaka, and Eneko Agirre. 2016 · 2016
Cited alongside, same era.
Enriching word vectors with subword information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017 · 2017
Cited alongside, same era.
An empirical analysis of nmt-derived interlingual embeddings and their use in parallel sentence identification
Cristina Espana-Bonet, Adám Csaba Varga, Alberto Barrón-Cedeno, and Josef van Genabith. 2017 · 2017
Cited alongside, same era.
Generating alignments using target foresight in attention-based neural machine translation
Jan-Thorsten Peter, Arne Nix, and Hermann Ney. 2017 · 2017
Cited alongside, same era.
Bilingual lexicon induction through unsupervised machine translation
Mikel Artetxe, Gorka Labaka, and Eneko Agirre. 2019 · 2019
Later among the works it cites.
BERT: Pre-training of deep bidirectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019 · 2019
Later among the works it cites.
Jointly learning to align and translate with transformer models
Sarthak Garg, Stephan Peitz, Udhyakumar Nallasamy, and Matthias Paulik. 2019 · 2019
Later among the works it cites.
Billion-scale similarity search with gpus
Jeff Johnson, Matthijs Douze, and Hervé Jégou. 2019 · 2019
Later among the works it cites.
Lost in evaluation: Misleading benchmarks for bilingual dictionary induction
Yova Kementchedjhieva, Mareike Hartmann, and Anders Søgaard. 2019 · 2019
Later among the works it cites.
The curious case of neural text degeneration
Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020 · 2020
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Offline bilingual word vectors, orthogonal transformations and the inverted softmax
Samuel L Smith, David HP Turban, Steven Hamblin, and Nils Y Hammerla. 2017 · 2017
Cited alongside, same era.
Adversarial training for unsupervised bilingual lexicon induction
Meng Zhang, Yang Liu, Huanbo Luan, and Maosong Sun. 2017 · 2017
Cited alongside, same era.
A robust self-learning method for fully unsupervised cross-lingual mappings of word embeddings
Mikel Artetxe, Gorka Labaka, and Eneko Agirre. 2018 · 2018
Cited alongside, same era.
Effective parallel corpus mining using bilingual sentence embeddings
Mandy Guo, Qinlan Shen, Yinfei Yang, Heming Ge, Daniel Cer, Gustavo Hernandez Abrego, Keith Stevens, Noah Constant, Yun-Hsuan Sung, Brian Strope, and Ray Kurzweil. 2018 · 2018
Cited alongside, same era.
Non-adversarial unsupervised word translation
Yedid Hoshen and Lior Wolf. 2018 · 2018
Cited alongside, same era.
Loss in translation: Learning bilingual word mapping with a retrieval criterion
Armand Joulin, Piotr Bojanowski, Tomas Mikolov, Herve Jegou, and Edouard Grave. 2018 · 2018
Cited alongside, same era.
Word translation without parallel data
Guillaume Lample, Alexis Conneau, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou. 2018 · 2018
Cited alongside, same era.
Later among the works it cites.
Unsupervised bitext mining and translation via self-trained contextual embeddings
Phillip Keung, Julian Salazar, Yichao Lu, and Noah A. Smith. 2020 · 2020
Later among the works it cites.
TALN/LS2N participation at the BUCC shared task: Bilingual dictionary induction from comparable corpora
Martin Laville, Amir Hazem, and Emmanuel Morin. 2020 · 2020
Later among the works it cites.
Overview of the fourth BUCC shared task: Bilingual dictionary induction from comparable corpora
Reinhard Rapp, Pierre Zweigenbaum, and Serge Sharoff. 2020 · 2020
Later among the works it cites.
Masoud Jalili Sabet, Philipp Dufter, and Hinrich Schutze. 2020 · 2020
Later among the works it cites.
LMU bilingual dictionary induction system with word surface similarity scores for BUCC 2020
Silvia Severini, Viktor Hangya, Alexander Fraser, and Hinrich Schütze. 2020 · 2020
Later among the works it cites.
Cross-lingual retrieval for iterative self-supervised training
Chau Tran, Yuqing Tang, Xian Li, and Jiatao Gu. 2020 · 2020
Later among the works it cites.