Fetching the paper…
Reading the bibliography…
We analyze three critical components of word embedding training: the model, the corpus, and the training parameters.
Distributional structure
Z. S. Harris · 1954
Earlier work this paper cites.
A solution to plato’s problem: The latent semantic analysis theory of acquisition, induction, and representation of knowledge
T. K. Landauer and S. T. Dumais · 1997
Earlier work this paper cites.
Early stopping-but when?
L. Prechelt · 1998
Earlier work this paper cites.
Placing search in context: The concept revisited
L. Finkelstein, E. Gabrilovich, Y. Matias, E. Rivlin, Z. Solan, G. Wolfman, and E. Ruppin · 2002
Earlier work this paper cites.
On the computational basis of learning and cognition: Arguments from lsa
T. K. Landauer · 2002
Earlier work this paper cites.
The computation of word associations: comparing syntagmatic and paradigmatic approaches
R. Rapp · 2002
Earlier work this paper cites.
A Neural Probabilistic Language Model
Y. Bengio, R. Ducharme, P. Vincent, and C. Jauvin · 2003
Earlier work this paper cites.
Feature-rich part-of-speech tagging with a cyclic dependency network
K. Toutanova, D. Klein, C. D. Manning, and Y. Singer · 2003
Earlier work this paper cites.
Hierarchical probabilistic neural network language model
F. Morin and Y. Bengio · 2005
Earlier work this paper cites.
The Word-Space Model
M. Sahlgren · 2006
Earlier work this paper cites.
Three new graphical models for statistical language modelling
A. Mnih and G. Hinton · 2007
Earlier work this paper cites.
A unified architecture for natural language processing: Deep neural networks with multitask learning
R. Collobert and J. Weston · 2008
Earlier work this paper cites.
A scalable hierarchical distributed language model
A. Mnih and G. E. Hinton · 2009
Cited alongside, same era.
Design challenges and misconceptions in named entity recognition
L. Ratinov and D. Roth · 2009
Cited alongside, same era.
Why does unsupervised pre-training help deep learning?
D. Erhan, Y. Bengio, A. Courville, P.-A. Manzagol, P. Vincent, and S. Bengio · 2010
Cited alongside, same era.
Word representations: a simple and general method for semi-supervised learning
J. Turian, L. Ratinov, and Y. Bengio · 2010
Cited alongside, same era.
From frequency to meaning: Vector space models of semantics
P. D. Turney, P. Pantel, et al · 2010
Cited alongside, same era.
Natural language processing (almost) from scratch
R. Collobert, J. Weston, L. Bottou, M. Karlen, K. Kavukcuoglu, and P. Kuksa · 2011
Cited alongside, same era.
Linguistic regularities in continuous space word representations
T. Mikolov, W.-t. Yih, and G. Zweig · 2013
Later among the works it cites.
Learning word embeddings efficiently with noise-contrastive estimation
A. Mnih and K. Kavukcuoglu · 2013
Later among the works it cites.
Recursive deep models for semantic compositionality over a sentiment treebank
R. Socher, A. Perelygin, J. Wu, J. Chuang, C. D. Manning, A. Ng, and C. Potts · 2013
Later among the works it cites.
Don’t count, predict! a systematic comparison of context-counting vs. context-predicting semantic vectors
M. Baroni, G. Dinu, and G. Kruszewski · 2014
Later among the works it cites.
Distributed Representations for Compositional Semantics
K. M. Hermann · 2014
Later among the works it cites.
Convolutional neural networks for sentence classification
Y. Kim · 2014
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Adaptive subgradient methods for online learning and stochastic optimization
J. Duchi, E. Hazan, and Y. Singer · 2011
Cited alongside, same era.
Learning word vectors for sentiment analysis
A. L. Maas, R. E. Daly, P. T. Pham, D. Huang, A. Y. Ng, and C. Potts · 2011
Cited alongside, same era.
Noise-contrastive estimation of unnormalized statistical models, with applications to natural image statistics
M. U. Gutmann and A. Hyvärinen · 2012
Cited alongside, same era.
Size (and domain) matters: Evaluating semantic word space representations for biomedical text
P. Stenetorp, H. Soyer, S. Pyysalo, S. Ananiadou, and T. Chikayama · 2012
Cited alongside, same era.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Cited alongside, same era.
Distributed representations of words and phrases and their compositionality
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Cited alongside, same era.
Later among the works it cites.
Word embeddings through hellinger pca
R. Lebret and R. Collobert · 2014
Later among the works it cites.
Neural word embedding as implicit matrix factorization
O. Levy and Y. Goldberg · 2014
Later among the works it cites.
Evaluating neural word representations in tensor-based compositional settings
D. Milajevs, D. Kartsaklis, M. Sadrzadeh, and M. Purver · 2014
Later among the works it cites.
GloVe : Global Vectors for Word Representation
J. Pennington, R. Socher, and C. D. Manning · 2014
Later among the works it cites.
Improving distributional similarity with lessons learned from word embeddings
O. Levy, Y. Goldberg, and I. Dagan · 2015
Closest in time.