Fetching the paper…
Reading the bibliography…
Recurrent neural networks (RNNs), particularly long short-term memory (LSTM), have gained much attention in automatic speech recognition (ASR).
D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning representations by back-propagating errors,” Nature , vol. 323, no. 6088, pp. 533–536, 1986, 10.1038/323533a0. [Online]. Available: http://dx.doi.org/10.1038/323533a0
1986
Earlier work this paper cites.
Y. Bengio, P. Simard, and P. Frasconi, “Learning long-term dependencies with gradient descent is difficult,” Neural Networks, IEEE Transactions on , vol. 5, no. 2, pp. 157–166, 1994
1994
Earlier work this paper cites.
S. Hochreiter and J. Schmidhuber, “Long short-term memory,” Neural computation , vol. 9, no. 8, pp. 1735–1780, 1997
1997
Earlier work this paper cites.
H. Jaeger and H. Haas, “Harnessing nonlinearity: Predicting chaotic systems and saving energy in wireless communication,” Science , vol. 304, no. 5667, pp. 78–80, 2004
2004
Earlier work this paper cites.
A. Graves and J. Schmidhuber, “Framewise phoneme classification with bidirectional lstm and other neural network architectures,” Neural Networks , vol. 18, no. 5, pp. 602–610, 2005
2005
Earlier work this paper cites.
C. Bucilu¨£, R. Caruana, and A. Niculescu-Mizil, “Model compression,” in Proceedings of the 12th ACM SIGKDD international conference on Knowledge discovery and data mining . ACM, 2006, pp. 535–541
2006
Earlier work this paper cites.
G. E. Hinton and R. R. Salakhutdinov, “Reducing the dimensionality of data with neural networks,” Science , vol. 313, no. 5786, pp. 504–507, 2006
2006
Earlier work this paper cites.
Y. Bengio, P. Lamblin, D. Popovici, H. Larochelle et al. , “Greedy layer-wise training of deep networks,” Advances in neural information processing systems , vol. 19, p. 153, 2007
2007
Earlier work this paper cites.
J. Martens, “Deep learning via hessian-free optimization,” in Proceedings of the 27th International Conference on Machine Learning (ICML-10) , 2010, pp. 735–742
2010
Cited alongside, same era.
D. Erhan, Y. Bengio, A. Courville, P.-A. Manzagol, P. Vincent, and S. Bengio, “Why does unsupervised pre-training help deep learning?” The Journal of Machine Learning Research , vol. 11, pp. 625–660, 2010
2010
Cited alongside, same era.
J. Martens and I. Sutskever, “Learning recurrent neural networks with hessian-free optimization,” in Proceedings of the 28th International Conference on Machine Learning (ICML-11) , 2011, pp. 1033–1040
2011
Cited alongside, same era.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, J. Silovsky, G. Stemmer, and K. Vesely, “The kaldi speech recognition toolkit,” in IEEE 2011 Workshop on Automatic Speech Recognition and Understanding . IEEE Signal Processing Society, Dec. 2011, iEEE Catalog No.: CFP11SRW-USB
A. Graves and N. Jaitly, “Towards end-to-end speech recognition with recurrent neural networks,” in Proceedings of the 31st International Conference on Machine Learning (ICML-14) , 2014, pp. 1764–1772
2014
Later among the works it cites.
H. Sak, A. Senior, and F. Beaufays, “Long short-term memory recurrent neural network architectures for large scale acoustic modeling,” in Proceedings of the Annual Conference of International Speech Communication Association (INTERSPEECH) , 2014
2014
Later among the works it cites.
J. Ba and R. Caruana, “Do deep nets really need to be deep?” in Advances in Neural Information Processing Systems , 2014, pp. 2654–2662
2014
Later among the works it cites.
G. E. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” in NIPS 2014 Deep Learning Workshop , 2014
2014
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
2011
Cited alongside, same era.
L. Deng and D. Yu, “Deep learning: Methods and applications,” Foundations and Trends in Signal Processing , vol. 7, no. 3-4, pp. 197–387, 2013. [Online]. Available: http://dx.doi.org/10.1561/2000000039
2013
Cited alongside, same era.
A. Graves, A.-R. Mohamed, and G. Hinton, “Speech recognition with deep recurrent neural networks,” in Proceedings of IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) . IEEE, 2013, pp. 6645–6649
2013
Cited alongside, same era.
I. Sutskever, J. Martens, G. Dahl, and G. Hinton, “On the importance of initialization and momentum in deep learning,” in Proceedings of the 30th International Conference on Machine Learning (ICML-13) , 2013, pp. 1139–1147
2013
Cited alongside, same era.
J. Li, R. Zhao, J.-T. Huang, and Y. Gong, “Learning small-size DNN with output-distribution-based criteria,” in Proceedings of the Annual Conference of International Speech Communication Association (INTERSPEECH) , 2014
2014
Later among the works it cites.
2014
Later among the works it cites.