Deciphering foreign language by combining language models and context vectors
Malte Nuhn, Arne Mauser, and Hermann Ney. 2012 · 2012
Cited alongside, same era.
Adam: A method for stochastic optimization
Original
Diederik P Kingma and Jimmy Ba. 2014 · 2014
Cited alongside, same era.
Bridging nonlinearities and stochastic regularizers with gaussian error linear units
Dan Hendrycks and Kevin Gimpel. 2016 · 2016
Cited alongside, same era.
Unsupervised neural machine translation
Original
Mikel Artetxe, Gorka Labaka, Eneko Agirre, and Kyunghyun Cho. 2017 · 2017
Cited alongside, same era.
Enriching word vectors with subword information
Piotr Bojanowski, Edouard Grave, Armand Joulin, and Tomas Mikolov. 2017 · 2017
Cited alongside, same era.
Word translation without parallel data
Original
Alexis Conneau, Guillaume Lample, Marc’Aurelio Ranzato, Ludovic Denoyer, and Hervé Jégou. 2017 · 2017
Cited alongside, same era.
Unsupervised machine translation using monolingual corpora only
Original
Guillaume Lample, Ludovic Denoyer, and Marc’Aurelio Ranzato. 2017 · 2017
Cited alongside, same era.
Automatic differentiation in pytorch
Adam Paszke, Sam Gross, Soumith Chintala, Gregory Chanan, Edward Yang, Zachary DeVito, Zeming Lin, Alban Desmaison, Luca Antiga, and Adam Lerer. 2017 · 2017
Cited alongside, same era.
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. 2017 · 2017
Cited alongside, same era.
A robust self-learning method for fully unsupervised cross-lingual mappings of word embeddings
Mikel Artetxe, Gorka Labaka, and Eneko Agirre. 2018a
Cited in the paper.
Improving neural machine translation models with monolingual data
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016a
Cited in the paper.
Neural machine translation of rare words with subword units
Rico Sennrich, Barry Haddow, and Alexandra Birch. 2016b
Cited in the paper.