Greedy layer-wise training of deep networks
Bengio, Yoshua, Lamblin, Pascal, Popovici, Dan, Larochelle, Hugo, et al · 2007
Cited alongside, same era.
Learning bilingual lexicons from monolingual corpora
Haghighi, Aria, Liang, Percy, Berg-Kirkpatrick, Taylor, and Klein, Dan · 2008
Cited alongside, same era.
Attacking decipherment problems optimally with low-order n-gram models
Ravi, Sujith and Knight, Kevin · 2008
Cited alongside, same era.
A statistical model for lost language decipherment
Snyder, Benjamin, Barzilay, Regina, and Knight, Kevin · 2010
Cited alongside, same era.
An analysis of single-layer networks in unsupervised feature learning
Coates, Adam, Ng, Andrew Y, and Lee, Honglak · 2011
Cited alongside, same era.
Unsupervised and transfer learning challenge: a deep learning approach
Mesnil, Grégoire, Dauphin, Yann, Glorot, Xavier, Rifai, Salah, Bengio, Yoshua, Goodfellow, Ian J, Lavoie, Erick, Muller, Xavier, Desjardins, Guillaume, Warde-Farley, David, et al · 2012
Cited alongside, same era.
Auto-encoding variational bayes
Original
Kingma, Diederik P and Welling, Max · 2013
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Original
Bahdanau, Dzmitry, Cho, Kyunghyun, and Bengio, Yoshua · 2014
Cited alongside, same era.
Efficient estimation of word representations in vector space
Original
Mikolov, Tomas, Chen, Kai, Corrado, Greg, and Dean, Jeffrey
Cited in the paper.
Exploiting similarities among languages for machine translation
Original
Mikolov, Tomas, Le, Quoc V, and Sutskever, Ilya
Cited in the paper.