Fetching the paper…
Reading the bibliography…
The recently introduced continuous Skip-gram model is an efficient method for learning high-quality distributed vector representations that capture a large number of precise syntactic and semantic word relationships.
Learning representations by back-propagating errors
David E Rumelhart, Geoffrey E Hintont, and Ronald J Williams · 1986
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Janvin · 2003
Earlier work this paper cites.
Hierarchical probabilistic neural network language model
Frederic Morin and Yoshua Bengio · 2005
Earlier work this paper cites.
Continuous space language models
Holger Schwenk · 2007
Earlier work this paper cites.
A unified architecture for natural language processing: deep neural networks with multitask learning
Ronan Collobert and Jason Weston · 2008
Earlier work this paper cites.
A scalable hierarchical distributed language model
Andriy Mnih and Geoffrey E Hinton · 2009
Earlier work this paper cites.
Word representations: a simple and general method for semi-supervised learning
Joseph Turian, Lev Ratinov, and Yoshua Bengio · 2010
Earlier work this paper cites.
From frequency to meaning: Vector space models of semantics
Peter D. Turney and Patrick Pantel · 2010
Cited alongside, same era.
Domain adaptation for large-scale sentiment classification: A deep learning approach
Xavier Glorot, Antoine Bordes, and Yoshua Bengio · 2011
Cited alongside, same era.
Extensions of recurrent neural network language model
Tomas Mikolov, Stefan Kombrink, Lukas Burget, Jan Cernocky, and Sanjeev Khudanpur · 2011
Cited alongside, same era.
Strategies for Training Large Scale Neural Network Language Models
Tomas Mikolov, Anoop Deoras, Daniel Povey, Lukas Burget and Jan Cernocky · 2011
Cited alongside, same era.
Parsing natural scenes and natural language with recursive neural networks
Richard Socher, Cliff C. Lin, Andrew Y. Ng, and Christopher D. Manning · 2011
Cited alongside, same era.
Wsabie: Scaling up to large vocabulary image annotation
Jason Weston, Samy Bengio, and Nicolas Usunier · 2011
Statistical Language Models Based on Neural Networks
Tomas Mikolov · 2012
Later among the works it cites.
A fast and simple algorithm for training neural probabilistic language models
Andriy Mnih and Yee Whye Teh · 2012
Later among the works it cites.
Semantic Compositionality Through Recursive Matrix-Vector Spaces
Richard Socher, Brody Huval, Christopher D. Manning, and Andrew Y. Ng · 2012
Later among the works it cites.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Closest in time.
Linguistic Regularities in Continuous Space Word Representations
Tomas Mikolov, Wen-tau Yih and Geoffrey Zweig · 2013
Closest in time.
Distributional semantics beyond words: Supervised learning of analogy and paraphrase
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
Noise-contrastive estimation of unnormalized statistical models, with applications to natural image statistics
Michael U Gutmann and Aapo Hyvärinen · 2012
Cited alongside, same era.
Peter D. Turney · 2013
Closest in time.