Fetching the paper…
Reading the bibliography…
Neural language models learn word representations, or embeddings, that capture rich linguistic and conceptual information.
A synopsis of linguistic theory 1930-1955 , pp. 1–32
Firth, J, R · 1957
Earlier work this paper cites.
A solution to plato’s problem: The latent semantic analysis theory of acquisition, induction, and representation of knowledge
Landauer, Thomas K and Dumais, Susan T · 1997
Earlier work this paper cites.
Quick training of probabilistic neural nets by importance sampling
Bengio, Yoshua and Sénécal, Jean-Sébastien · 2003
Earlier work this paper cites.
A neural probabilistic language model
Bengio, Yoshua, Ducharme, Réjean, Vincent, Pascal, and Janvin, Christian · 2003
Earlier work this paper cites.
The university of south florida free association, rhyme, and word fragment norms
Nelson, Douglas L, McEvoy, Cathy L, and Schreiber, Thomas A · 2004
Earlier work this paper cites.
Hierarchical probabilistic neural network language model
Morin, Frederic and Bengio, Yoshua · 2005
Earlier work this paper cites.
A unified architecture for natural language processing: Deep neural networks with multitask learning
Collobert, Ronan and Weston, Jason · 2008
Earlier work this paper cites.
Learning bilingual lexicons from monolingual corpora
Haghighi, Aria, Liang, Percy, Berg-Kirkpatrick, Taylor, and Klein, Dan · 2008
Earlier work this paper cites.
A study on similarity and relatedness using distributional and wordnet-based approaches
Agirre, Eneko, Alfonseca, Enrique, Hall, Keith, Kravalova, Jana, Pasca, Marius, and Soroa, Aitor · 2009
Earlier work this paper cites.
A scalable hierarchical distributed language model
Mnih, Andriy and Hinton, Geoffrey E · 2009
Earlier work this paper cites.
Theano: a CPU and GPU math expression compiler
Bergstra, James, Breuleux, Olivier, Bastien, Frédéric, Lamblin, Pascal, Pascanu, Razvan, Desjardins, Guillaume, Turian, Joseph, Warde-Farley, David, and Bengio, Yoshua · 2010
Earlier work this paper cites.
From frequency to meaning: Vector space models of semantics
Turney, Peter D and Pantel, Patrick · 2010
Cited alongside, same era.
Identifying word translations from comparable corpora using latent topic models
Vulić, Ivan, De Smet, Wim, and Moens, Marie-Francine · 2011
Cited alongside, same era.
Inducing crosslingual distributed representations of words
A. Klementiev, I. Titov and Bhattarai, B · 2012
Cited alongside, same era.
Theano: new features and speed improvements
Bastien, Frédéric, Lamblin, Pascal, Pascanu, Razvan, Bergstra, James, Goodfellow, Ian J., Bergeron, Arnaud, Bouchard, Nicolas, and Bengio, Yoshua · 2012
Cited alongside, same era.
Recurrent continuous translation models
Kalchbrenner, Nal and Blunsom, Phil · 2013
Cited alongside, same era.
Neural machine translation by jointly learning to align and translate
Improving vector space word representations using multilingual correlation
Faruqui, Manaal and Dyer, Chris · 2014
Closest in time.
Multilingual Distributed Representations without Word Alignment
Hermann, Karl Moritz and Blunsom, Phil · 2014
Closest in time.
Learning abstract concepts from multi-modal data: Since you probably can’t see what i mean
Hill, Felix and Korhonen, Anna · 2014
Closest in time.
Simlex-999: Evaluating semantic models with (genuine) similarity estimation
Hill, Felix, Reichart, Roi, and Korhonen, Anna · 2014
Closest in time.
On using very large target vocabulary for neural machine translation
Jean, Sébastien, Cho, Kyunghyun, Memisevic, Roland, and Bengio, Yoshua · 2014
Closest in time.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Bahdanau, Dzmitry, Cho, Kyunghyun, and Bengio, Yoshua · 2014
Cited alongside, same era.
Don’t count, predict! a systematic comparison of context-counting vs. context-predicting semantic vectors
Baroni, Marco, Dinu, Georgiana, and Kruszewski, Germán · 2014
Cited alongside, same era.
Multimodal distributional semantics
Bruni, Elia, Tran, Nam-Khanh, and Baroni, Marco · 2014
Cited alongside, same era.
An Autoencoder Approach to Learning Bilingual Word Representations
Chandar, Sarath, Lauly, Stanislas, Larochelle, Hugo, Khapra, Mitesh M., Ravindran, Balaraman, Raykar, Vikas, and Saha, Amrita · 2014
Cited alongside, same era.
Learning phrase representations using RNN encoder-decoder for statistical machine translation
Cho, Kyunghyun, van Merrienboer, Bart, Gulcehre, Caglar, Bougares, Fethi, Schwenk, Holger, and Bengio, Yoshua · 2014
Cited alongside, same era.
Fast and robust neural network joint models for statistical machine translation
Devlin, Jacob, Zbib, Rabih, Huang, Zhongqiang, Lamar, Thomas, Schwartz, Richard, and Makhoul, John · 2014
Cited alongside, same era.
Inducing crosslingual distributed representations of words
Klementiev, Alexandre, Titov, Ivan, and Bhattarai, Binod
Cited in the paper.
Learning Bilingual Word Representations by Marginalizing Alignments
Kočiský, Tomáš, Hermann, Karl Moritz, and Blunsom, Phil · 2014
Closest in time.
Dependency-based word embeddings
Levy, Omer and Goldberg, Yoav · 2014
Closest in time.
Addressing the rare word problem in neural machine translation
Luong, Thang, Sutskever, Ilya, Le, Quoc V, Vinyals, Oriol, and Zaremba, Wojciech · 2014
Closest in time.
Glove: Global vectors for word representation
Pennington, Jeffrey, Socher, Richard, and Manning, Christopher · 2014
Closest in time.
Sequence to sequence learning with neural networks
Sutskever, Ilya, Vinyals, Oriol, and Le, Quoc V · 2014
Closest in time.