Fetching the paper…
Reading the bibliography…
Word embedding models such as the skip-gram learn vector representations of words' semantic relationships, and document embedding models learn similar representations for documents.
Learning distributed representations of concepts
Geoffrey E Hinton et al · 1986
Earlier work this paper cites.
Training products of experts by minimizing contrastive divergence
Geoffrey E Hinton · 2002
Earlier work this paper cites.
A neural probabilistic language model
Yoshua Bengio, Réjean Ducharme, Pascal Vincent, and Christian Jauvin · 2003
Earlier work this paper cites.
Latent Dirichlet allocation
David M Blei, Andrew Y Ng, and Michael I Jordan · 2003
Earlier work this paper cites.
The author-topic model for authors and documents
Michal Rosen-Zvi, Thomas Griffiths, Mark Steyvers, and Padhraic Smyth · 2004
Earlier work this paper cites.
Model compression
Cristian Buciluǎ, Rich Caruana, and Alexandru Niculescu-Mizil · 2006
Earlier work this paper cites.
Topics in semantic representation
Thomas L Griffiths, Mark Steyvers, and Joshua B Tenenbaum · 2007
Earlier work this paper cites.
The distributional hypothesis
Magnus Sahlgren · 2008
Earlier work this paper cites.
Noise-contrastive estimation: A new estimation principle for unnormalized statistical models
Michael Gutmann and Aapo Hyvärinen · 2010
Earlier work this paper cites.
Natural language processing (almost) from scratch
Ronan Collobert, Jason Weston, Léon Bottou, Michael Karlen, Koray Kavukcuoglu, and Pavel Kuksa · 2011
Cited alongside, same era.
Learning word vectors for sentiment analysis
Andrew L Maas, Raymond E Daly, Peter T Pham, Dan Huang, Andrew Y Ng, and Christopher Potts · 2011
Cited alongside, same era.
Optimizing semantic coherence in topic models
David Mimno, Hanna M Wallach, Edmund Talley, Miriam Leenders, and Andrew McCallum · 2011
Cited alongside, same era.
Noise-contrastive estimation of unnormalized statistical models, with applications to natural image statistics
Michael U Gutmann and Aapo Hyvärinen · 2012
Cited alongside, same era.
Efficient estimation of word representations in vector space
Tomas Mikolov, Kai Chen, Greg Corrado, and Jeffrey Dean · 2013
Cited alongside, same era.
Distributed representations of sentences and documents
Quoc Le and Tomas Mikolov · 2014
Later among the works it cites.
Neural word embedding as implicit matrix factorization
Omer Levy and Yoav Goldberg · 2014
Later among the works it cites.
Reducing the sampling complexity of topic models
Aaron Q Li, Amr Ahmed, Sujith Ravi, and Alexander J Smola · 2014
Later among the works it cites.
GloVe: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning · 2014
Later among the works it cites.
Gaussian LDA for topic models with word embeddings
Rajarshi Das, Manzil Zaheer, and Chris Dyer · 2015
Later among the works it cites.
Distilling the knowledge in a neural network
Geoffrey Hinton, Oriol Vinyals, and Jeff Dean · 2015
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean · 2013
Cited alongside, same era.
Learning word embeddings efficiently with noise-contrastive estimation
Andriy Mnih and Koray Kavukcuoglu · 2013
Cited alongside, same era.
Decoding with large-scale neural language models improves translation
Ashish Vaswani, Yinggong Zhao, Victoria Fossum, and David Chiang · 2013
Cited alongside, same era.
Later among the works it cites.
Topical word embeddings
Yang Liu, Zhiyuan Liu, Tat-Seng Chua, and Maosong Sun · 2015
Later among the works it cites.
Mixed membership word embeddings for computational social science
J. R. Foulds · 2018
Later among the works it cites.