A unified architecture for natural language processing: Deep neural networks with multitask learning
Ronan Collobert and Jason Weston. 2008 · 2008
Cited alongside, same era.
Neural word embedding as implicit matrix factorization
Omer Levy and Yoav Goldberg. 2014 · 2014
Cited alongside, same era.
GloVe: Global vectors for word representation
Jeffrey Pennington, Richard Socher, and Christopher Manning. 2014 · 2014
Cited alongside, same era.
Zipf’s word frequency law in natural language: A critical review and future directions
Steven T Piantadosi. 2014 · 2014
Cited alongside, same era.
A latent variable model approach to PMI-based word embeddings
Sanjeev Arora, Yuanzhi Li, Yingyu Liang, Tengyu Ma, and Andrej Risteski. 2016 · 2016
Cited alongside, same era.
A simple but tough-to-beat baseline for sentence embeddings
Sanjeev Arora, Yingyu Liang, and Tengyu Ma. 2017 · 2017
Cited alongside, same era.
Efficient estimation of word representations in vector space
Original
Tomas Mikolov, Kai Chen, Greg S Corrado, and Jeff Dean. 2013a
Cited in the paper.
Distributed representations of words and phrases and their compositionality
Tomas Mikolov, Ilya Sutskever, Kai Chen, Greg S Corrado, and Jeff Dean. 2013b
Cited in the paper.