Fetching the paper…
Reading the bibliography…
This paper develops a model that addresses sentence embedding, a hot topic in current natural language processing research, using recurrent neural networks with Long Short-Term Memory (LSTM) cells.
Y. Nesterov, “A method of solving a convex programming problem with convergence rate o (1/k2),” Soviet Mathematics Doklady , vol. 27, pp. 372–376, 1983
1983
Earlier work this paper cites.
J. L. Elman, “Finding structure in time,” Cognitive Science , vol. 14, no. 2, pp. 179–211, 1990
1990
Earlier work this paper cites.
A. J. Robinson, “An application of recurrent nets to phone probability estimation,” IEEE Transactions on Neural Networks , vol. 5, no. 2, pp. 298–305, August 1994
1994
Earlier work this paper cites.
L. Deng, K. Hassanein, and M. Elmasry, “Analysis of the correlation structure for a neural predictive model with application to speech recognition,” Neural Networks , vol. 7, no. 2, pp. 331–339, 1994
1994
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural Comput. , vol. 9, no. 8, pp. 1735–1780, Nov. 1997
1997
Earlier work this paper cites.
F. A. Gers, J. Schmidhuber, and F. Cummins, “Learning to forget: Continual prediction with lstm,” Neural Computation , vol. 12, pp. 2451–2471, 1999
1999
Earlier work this paper cites.
T. Hofmann, “Probabilistic latent semantic analysis,” in In Proc. of Uncertainty in Artificial Intelligence, UAI’99 , 1999, pp. 289–296
1999
Earlier work this paper cites.
K. Järvelin and J. Kekäläinen, “Ir evaluation methods for retrieving highly relevant documents,” in Proceedings of the 23rd Annual International ACM SIGIR Conference on Research and Development in Information Retrieval , ser. SIGIR. ACM, 2000, pp. 41–48
2000
Earlier work this paper cites.
F. A. Gers, N. N. Schraudolph, and J. Schmidhuber, “Learning precise timing with lstm recurrent networks,” J. Mach. Learn. Res. , vol. 3, pp. 115–143, Mar. 2003
2003
Earlier work this paper cites.
R. Collobert and J. Weston, “A unified architecture for natural language processing: Deep neural networks with multitask learning,” in International Conference on Machine Learning, ICML , 2008
2008
Earlier work this paper cites.
J. Gao, W. Yuan, X. Li, K. Deng, and J.-Y. Nie, “Smoothing clickthrough data for web search ranking,” in Proceedings of the 32Nd International ACM SIGIR Conference on Research and Development in Information Retrieval , ser. SIGIR ’09. New York, NY, USA: ACM, 2009, pp. 355–362
2009
Earlier work this paper cites.
T. Mikolov, M. Karafiát, L. Burget, J. Cernockỳ, and S. Khudanpur, “Recurrent neural network based language model.” in Proc. INTERSPEECH , Makuhari, Japan, September 2010, pp. 1045–1048
2010
Earlier work this paper cites.
R. Řehůřek and P. Sojka, “Software Framework for Topic Modelling with Large Corpora,” in Proceedings of the LREC 2010 Workshop on New Challenges for NLP Frameworks . Valletta, Malta: ELRA, May 2010, pp. 45–50, http://is.muni.cz/publication/884893/en
2010
Earlier work this paper cites.
R. Socher, J. Pennington, E. H. Huang, A. Y. Ng, and C. D. Manning, “Semi-supervised recursive autoencoders for predicting sentiment distributions,” in Proceedings of the Conference on Empirical Methods in Natural Language Processing , ser. EMNLP ’11, 2011, pp. 151–161
2011
Earlier work this paper cites.
J. Gao, K. Toutanova, and W.-t. Yih, “Clickthrough-based latent semantic models for web search,” ser. SIGIR ’11. ACM, 2011, pp. 675–684
2011
Earlier work this paper cites.
G. E. Dahl, D. Yu, L. Deng, and A. Acero, “Large vocabulary continuous speech recognition with context-dependent DBN-HMMs,” in Proc. IEEE ICASSP , Prague, Czech, May 2011, pp. 4688–4691
2011
Earlier work this paper cites.
D. Yu and L. Deng, “Deep learning and its applications to signal and information processing [exploratory dsp],” IEEE Signal Processing Magazine , vol. 28, no. 1, pp. 145 –154, jan. 2011
2011
Cited alongside, same era.
A. Graves, “Sequence transduction with recurrent neural networks,” in Representation Learning Workshp, ICML , 2012
2012
Cited alongside, same era.
G. Dahl, D. Yu, L. Deng, and A. Acero, “Context-dependent pre-trained deep neural networks for large-vocabulary speech recognition,” Audio, Speech, and Language Processing, IEEE Transactions on , vol. 20, no. 1, pp. 30 –42, jan. 2012
2012
Cited alongside, same era.
G. Hinton, L. Deng, D. Yu, G. E. Dahl, A. Mohamed, N. Jaitly, A. Senior, V. Vanhoucke, P. Nguyen, T. N. Sainath, and B. Kingsbury, “Deep neural networks for acoustic modeling in speech recognition: The shared views of four research groups,” IEEE Signal Processing Magazine , vol. 29, no. 6, pp. 82–97, November 2012
2012
Cited alongside, same era.
J. Chung, C. Gulcehre, K. Cho, and Y. Bengio, “Empirical evaluation of gated recurrent neural networks on sequence modeling,” NIPS Deep Learning Workshop , 2014
2014
Later among the works it cites.
Y. Shen, X. He, J. Gao, L. Deng, and G. Mesnil, “A latent semantic model with convolutional-pooling structure for information retrieval.” CIKM, November 2014
2014
Later among the works it cites.
N. Kalchbrenner, E. Grefenstette, and P. Blunsom, “A convolutional neural network for modelling sentences,” Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics , June 2014
2014
Later among the works it cites.
J. Zhang, S. Liu, M. Li, M. Zhou, and C. Zong, “Bilingually-constrained phrase embeddings for machine translation,” in Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (ACL) (Volume 1: Long Papers) , Baltimore, Maryland, 2014, pp. 111–121
2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
L. Deng, D. Yu, and J. Platt, “Scalable stacking and learning for building deep architectures,” in Proc. ICASSP , march 2012, pp. 2133 –2136
2012
Cited alongside, same era.
P.-S. Huang, X. He, J. Gao, L. Deng, A. Acero, and L. Heck, “Learning deep structured semantic models for web search using clickthrough data,” in Proceedings of the 22Nd ACM International Conference on Conference on Information & Knowledge Management , ser. CIKM ’13. ACM, 2013, pp. 2333–2338
2013
Cited alongside, same era.
T. Mikolov, I. Sutskever, K. Chen, G. S. Corrado, and J. Dean, “Distributed representations of words and phrases and their compositionality,” in Proceedings of Advances in Neural Information Processing Systems , 2013, pp. 3111–3119
2013
Cited alongside, same era.
2013
Cited alongside, same era.
A. Graves, A. Mohamed, and G. Hinton, “Speech recognition with deep recurrent neural networks,” in Proc. ICASSP , Vancouver, Canada, May 2013
2013
Cited alongside, same era.
Y. Bengio, N. Boulanger-Lewandowski, and R. Pascanu, “Advances in optimizing recurrent networks,” in Proc. ICASSP , Vancouver, Canada, May 2013
2013
Cited alongside, same era.
G. Mesnil, X. He, L. Deng, and Y. Bengio, “Investigation of recurrent-neural-network architectures and learning methods for spoken language understanding,” in Proc. INTERSPEECH , Lyon, France, August 2013
2013
Cited alongside, same era.
I. Sutskever, J. Martens, G. E. Dahl, and G. E. Hinton, “On the importance of initialization and momentum in deep learning,” in ICML (3)’13 , 2013, pp. 1139–1147
2013
Cited alongside, same era.
H. Sak, A. Senior, and F. Beaufays, “Long short-term memory recurrent neural network architectures for large scale acoustic modeling,” in Proceedings of the Annual Conference of International Speech Communication Association (INTERSPEECH) , 2014
2014
Later among the works it cites.
2014
Later among the works it cites.
J. Chen and L. Deng, “A primal-dual method for training recurrent neural networks constrained by the echo-state property,” in Proceedings of the International Conf. on Learning Representations (ICLR) , 2014
2014
Later among the works it cites.
L. Deng and J. Chen, “Sequence classification using high-level features extracted from deep neural networks,” in Proc. ICASSP , 2014
2014
Later among the works it cites.
J. Gao, P. Pantel, M. Gamon, X. He, L. Deng, and Y. Shen, “Modeling interestingness with deep neural networks,” in Proc. EMNLP , 2014
2014
Later among the works it cites.
J. Gao, X. He, W. tau Yih, and L. Deng, “Learning continuous phrase representations for translation modeling,” in Proc. ACL , 2014
2014
Later among the works it cites.
R. Kiros, Y. Zhu, R. Salakhutdinov, R. S. Zemel, A. Torralba, R. Urtasun, and S. Fidler, “Skip-thought vectors,” Advances in Neural Information Processing Systems (NIPS) , 2015
2015
Closest in time.
2015
Closest in time.
2015
Closest in time.
2015
Closest in time.
B. Hu, Z. Lu, H. Li, and Q. Chen, “Convolutional neural network architectures for matching natural language sentences,” in Advances in Neural Information Processing Systems 27 , 2014, pp. 2042–2050
2050
Closest in time.