Fetching the paper…
Reading the bibliography…
Word2Vec is a widely used algorithm for extracting low-dimensional vector representations of words.
G. A. Miller and W. G. Charles, “Contextual correlates of semantic similarity,” in Language and cognitive processes , 1991
1991
Earlier work this paper cites.
C. D. Manning and H. Schütze, Foundations of Statistical Natural Language Processing . Cambridge, MA, USA: MIT Press, 1999
1999
Earlier work this paper cites.
L. S. Blackford, J. Demmel, J. Dongarra, I. Duff, S. Hammarling, G. Henry, M. Heroux, L. Kaufman, A. Lumsdaine, A. Petitet, R. Pozo, K. Remington, and R. C. Whaley, “An updated set of basic linear algebra subprograms (blas),” ACM Trans. Mathematical Software , vol. 28, no. 2, pp. 135–151, 2002
2002
Earlier work this paper cites.
L. Finkelstein, E. Gabrilovich, Y. Matias, E. Rivlin, Z. Solan, G. Wolfman, and E. Ruppin, “Placing search in context: The concept revisited,” ACM Transactions on Information Systems , vol. 20, pp. 116–131, 2002
2002
Earlier work this paper cites.
R. Collobert and J. Weston, “A unified architecture for natural language processing: deep neural networks with multitask learning,” in Proceedings of the 25th international conference on Machine learning , 2008, pp. 160–167
2008
Earlier work this paper cites.
X. Glorot, A. Bordes, and Y. Bengio, “A unified architecture for natural language processing: deep neural networks with multitask learning,” in Proceedings of the 25th international conference on Machine learning , 2011, pp. 513–520
2011
Earlier work this paper cites.
F. Niu, B. Recht, C. Re, and S. J. Wright, “Hogwild: A lock-free approach to parallelizing stochastic gradient descent,” in Advances in Neural Information Processing Systems , 2011, pp. 693–701
2011
Cited alongside, same era.
J. Duchi, E. Hazan, and Y. Singer, “Adaptive subgradient methods for online learning and stochastic optimization,” Journal of Machine Learning Research , vol. 12, pp. 2121–2159, 2011
2011
Cited alongside, same era.
G. Hinton, “Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude,” 2012, cOURSERA: Neural Networks for Machine Learning
2012
Cited alongside, same era.
P. D. Turney, “Distributional semantics beyond words: Supervised learning of analogy and paraphrase,” in Transactions of the Association for Computational Linguistics (TACL) , 2013, pp. 353–366
2013
Cited alongside, same era.
K. Cho, B. van Merrienboer, C. Gulcehre, D. Bahdanau, F. Bougares, H. Schwenk, and Y. Bengio, “Learning phrase representations using rnn encoder-decoder for statistical machine translation,” in Proceedings of the 2002 Conference on Empirical Methods in Natural Language Processing (EMNLP) , 2014
2014
Later among the works it cites.
C. Chelba, T. Mikolov, M. Schuster, Q. Ge, T. Brants, P. Koehn, and T. Robinson, “One billion word benchmark for measuring progress in statistical language modeling,” in INTERSPEECH , 2014, pp. 2635–2639
2014
Later among the works it cites.
2014
Later among the works it cites.
J. Weston, S. Chopra, and A. Bordes, “Memory networks,” in International Conference on Learning Representations (ICLR) , 2015
2015
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean, “Distributed representations of words and phrases and their compositionality,” in Advances in Neural Information Processing Systems 26 , 2013, pp. 3111–3119
2013
Cited alongside, same era.
T. Mikolov, K. Chen, G. Corrado, and J. Dean, “Efficient estimation of word representations in vector space,” Proceedings of Workshop at ICLR , 2013
2013
Cited alongside, same era.
J. Canny, H. Zhao, Y. Chen, B. Jaros, and J. Mao, “Machine learning at the limit,” in IEEE International Conference on Big Data , 2015
2015
Later among the works it cites.