Fetching the paper…
Reading the bibliography…
We revisit skip-gram negative sampling (SGNS), one of the most popular neural-network based approaches to learning distributed word representation.
A stochastic approximation method
H. Robbins and S. Monro · 1951
Earlier work this paper cites.
Distributional structure
Z. Harris · 1954
Earlier work this paper cites.
Matrix analysis
R. Horn and C. Johnson · 1990
Earlier work this paper cites.
Numerical linear algebra
L. N. Trefethen and D. Bau III · 1997
Earlier work this paper cites.
A neural probabilistic language model
Y. Bengio, R. Ducharme, P. Vincent, and C. Jauvin · 2003
Earlier work this paper cites.
Hierarchical probabilistic neural network language model
F. Morin and Y. Bengio · 2005
Earlier work this paper cites.
Neural probabilistic language models
Y. Bengio, H. Schwenk, J.-S. Senécal, F. Morin, and J.-L. Gauvain · 2006
Earlier work this paper cites.
A unified architecture for natural language processing: Deep neural networks with multitask learning
R. Collobert and J. Weston · 2008
Earlier work this paper cites.
A scalable hierarchical distributed language model
A. Mnih and G. Hinton · 2009
Earlier work this paper cites.
Tensor decompositions and applications
T. G. Kolda and B. W. Bader · 2009
Earlier work this paper cites.
Recurrent neural network based language model
T. Mikolov, M. Karafiát, L. Burget, J. Černockỳ, and S. Khudanpur · 2010
Earlier work this paper cites.
Natural language processing (almost) from scratch
R. Collobert, J. Weston, L. Bottou, M. Karlen, K. Kavukcuoglu, and P. Kuksa · 2011
Earlier work this paper cites.
Structured output layer neural network language model
H. Le, I. Oparin, A. Allauzen, J.-L. Gauvain, and F. Yvon · 2011
Earlier work this paper cites.
Incremental gradient, subgradient, and proximal methods for convex optimization: A survey
D. P. Bertsekas · 2011
Earlier work this paper cites.
Parsing with compositional vector grammars
R. Socher, J. Bauer, C. Manning, and A. Ng · 2013
Earlier work this paper cites.
Linguistic regularities in continuous space word representations
T. Mikolov, W. Yih, and G. Zweig · 2013
Earlier work this paper cites.
Efficient estimation of word representations in vector space
T. Mikolov, K. Chen, G. Corrado, and J. Dean · 2013
Earlier work this paper cites.
Distributed representations of words and phrases and their compositionality
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean · 2013
Cited alongside, same era.
Convolutional neural networks for sentence classification
Y. Kim · 2014
Cited alongside, same era.
A neural network for factoid question answering over paragraphs
M. Iyyer, J. Boyd-Graber, L. Claudino, R. Socher, and H. Daumé III · 2014
Cited alongside, same era.
Don’t count, predict! a systematic comparison of context-counting vs. context-predicting semantic vectors
M. Baroni, G. Dinu, and G. Kruszewski · 2014
Cited alongside, same era.
GloVe: Global vectors for word representation
J. Pennington, R. Socher, and C. Manning · 2014
Cited alongside, same era.
word2vec explained: deriving Mikolov et al.’s negative-sampling word-embedding method
Where to look: Focus regions for visual question answering
K. Shih, S. Singh, and D. Hoiem · 2016
Later among the works it cites.
Neural architectures for named entity recognition
G. Lample, M. Ballesteros, S. Subramanian, K. Kawakami, and C. Dyer · 2016
Later among the works it cites.
Sparse word embeddings using ℓ 1 \ell_{1} regularized online learning
F. Sun, J. Guo, Y. Lan, J. Xu, and X. Cheng · 2016
Later among the works it cites.
Generalized low rank models
M. Udell, C. Horn, R. Zadeh, and S. Boyd · 2016
Later among the works it cites.
Stochastic variance reduction for nonconvex optimization
S. J. Reddi, A. Hefny, S. Sra, B. Poczos, and A. Smola · 2016
Later among the works it cites.
Advances in pre-training distributed word representations
T. Mikolov, E. Grave, P. Bojanowski, C. Puhrsch, and A. Joulin · 2017
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Y. Goldberg and O. Levy · 2014
Cited alongside, same era.
Neural word embedding as implicit matrix factorization
O. Levy and Y. Goldberg · 2014
Cited alongside, same era.
Logistic matrix factorization for implicit feedback data
C. Johnson · 2014
Cited alongside, same era.
Linguistic regularities in sparse and explicit word representations
O. Levy and Y. Goldberg · 2014
Cited alongside, same era.
Tensor decompositions for learning latent variable models
A. Anandkumar, R. Ge, D. Hsu, S. M. Kakade, and M. Telgarsky · 2014
Cited alongside, same era.
Context-and content-aware embeddings for query rewriting in sponsored search
M. Grbovic, N. Djuric, V. Radosavljevic, F. Silvestri, and N. Bhamidipati · 2015
Cited alongside, same era.
Adapting word2vec to named entity recognition
S. Sienčnik · 2015
Cited alongside, same era.
Later among the works it cites.
A simple regularization-based algorithm for learning cross-domain word embeddings
W. Yang, W. Lu, and V. Zheng · 2017
Later among the works it cites.
word2vec skip-gram with negative sampling is a weighted logistic PCA
A. J. Landgraf and J. Bellay · 2017
Later among the works it cites.
Stochastic quasi-newton methods for nonconvex stochastic optimization
X. Wang, S. Ma, D. Goldfarb, and W. Liu · 2017
Later among the works it cites.
Using negative curvature in solving nonlinear programs
D. Goldfarb, C. Mu, J. Wright, and C. Zhou · 2017
Later among the works it cites.
Riemannian optimization for skip-gram negative sampling
A. Fonarev, O. Grinchuk, G. Gusev, P. Serdyukov, and I. Oseledets · 2017
Later among the works it cites.
Greedy approaches to symmetric orthogonal tensor decomposition
C. Mu, D. Hsu, and D. Goldfarb · 2017
Later among the works it cites.
Word embeddings via tensor factorization
E. Bailey and S. Aeron · 2017
Later among the works it cites.
On the convergence of adam and beyond
S. J. Reddi, S. Kale, and S. Kumar · 2018
Closest in time.
Stochastic optimization using a trust-region method and random models
R. Chen, M. Menickelly, and K. Scheinberg · 2018
Closest in time.
Understanding composition of word embeddings via tensor decomposition
A. Frandsen and R. Ge · 2019
Closest in time.