Fetching the paper…
Reading the bibliography…
Deep neural networks (DNNs) are now a central component of nearly all state-of-the-art speech recognition systems.
Y. Nesterov, “A method of solving a convex programming problem with convergence rate o (1/k2),” in Soviet Mathematics Doklady , vol. 27, no. 2, 1983, pp. 372–376
1983
Earlier work this paper cites.
J. L. McClelland and J. L. Elman, “The trace model of speech perception,” Cognitive psychology , vol. 18, no. 1, pp. 1–86, 1986
1986
Earlier work this paper cites.
L. B. Bahl, P. de Souza, and R. P. Mercer, “Maximum mutual information estimation of hidden markov model parameters for speech recognition,” in ICASSP . IEEE, 1986
1986
Earlier work this paper cites.
D. Plaut, “Experiments on learning by back propagation.” 1986
1986
Earlier work this paper cites.
D. E. Rumelhart, G. E. Hinton, and R. J. Williams, “Learning internal representations by error propagation,” Parallel distributed processing: explorations in the microstructures of cognition, volume 2: psychological and biological models , vol. 76, p. 1555, 1986
1986
Earlier work this paper cites.
A. Waibel, T. Hanazawa, G. Hinton, K. Shikano, and K. J. Lang, “Phoneme recognition using time-delay neural networks,” Acoustics, Speech and Signal Processing, IEEE Transactions on , vol. 37, no. 3, pp. 328–339, 1989
1989
Earlier work this paper cites.
T. Robinson and F. Fallside, “A recurrent error propagation network speech recognition system,” Computer Speech & Language , vol. 5, no. 3, pp. 259–274, 1991
1991
Earlier work this paper cites.
H. Bourlard and N. Morgan, Connectionist Speech Recognition: A Hybrid Approach . Norwell, MA: Kluwer Academic Publishers, 1993
1993
Earlier work this paper cites.
S. Renals, N. Morgan, H. Bourlard, M. Cohen, and H. Franco, “Connectionist probability estimators in hmm speech recognition,” IEEE Transactions on Speech and Audio Processing , vol. 2, no. 1, pp. 161–174, 1994
1994
Earlier work this paper cites.
C. S. Lindsey and T. Lindblad, “Survey of neural network hardware,” in SPIE Symposium on OE/Aerospace Sensing and Dual Use Photonics . International Society for Optics and Photonics, 1995, pp. 1194–1205
1995
Earlier work this paper cites.
S. Young, G. Evermann, M. Gales, T. Hain, D. Kershaw, X. Liu, G. Moore, J. Odell, D. Ollason, D. Povey et al. , The HTK book . Entropic Cambridge Research Laboratory Cambridge, 1997, vol. 2
1997
Earlier work this paper cites.
V. Valtchev, J. Odell, P. C. Woodland, and S. J. Young, “Mmie training of large vocabulary recognition systems,” Speech Communication , vol. 22, no. 4, pp. 303–314, 1997
1997
Earlier work this paper cites.
Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based Learning Applied to Document Recognition,” Proceedings of the IEEE , vol. 86, no. 11, pp. 2278–2324, 1998
1998
Earlier work this paper cites.
D. Ellis and N. Morgan, “Size matters: An empirical study of neural network training for large vocabulary continuous speech recognition,” in ICASSP . IEEE, 1999, pp. 1013–1016
1999
Earlier work this paper cites.
H. Hermansky, D. Ellis, and S. Sharma, “Tandem connectionist feature extraction for conventional hmm systems,” in ICASSP , vol. 3. IEEE, 2000, pp. 1635–1638
2000
Earlier work this paper cites.
D. Jurafsky and J. H. Martin, Speech and language processing: An introduction to natural language processing, computational linguistics, and speech recognition . Prentice Hall, 2000
2000
Earlier work this paper cites.
J. Kaiser, B. Horvat, and Z. Kacic, “A novel loss function for the overall risk criterion based discriminative training of hmm models,” in ICSLP , 2000
2000
Earlier work this paper cites.
R. Caruana, S. Lawrence, and L. Giles, “Overfitting in Neural Nets: Backpropagation, Conjugate Gradient, and Early Stopping,” in NIPS , 2000
2000
Earlier work this paper cites.
K. Oh and K. Jung, “Gpu implementation of neural networks,” Pattern Recognition , vol. 37, no. 6, pp. 1311–1314, 2004
2004
Earlier work this paper cites.
C. Cieri, D. Miller, and K. Walker, “The fisher corpus: a resource for the next generations of speech-to-text.” in LREC , vol. 4, 2004, pp. 69–71
2004
Earlier work this paper cites.
Z. Luo, H. Liu, and X. Wu, “Artificial neural network computation on graphic process unit,” in IJCNN . IEEE, 2005, pp. 622–626
2005
Earlier work this paper cites.
G. Hinton, S. Osindero, and Y. W. Teh, “A fast learning algorithm for deep belief nets,” Neural computation , vol. 18, no. 7, pp. 1527–1554, 2006
2006
Earlier work this paper cites.
M. Gales and S. Young, “The application of hidden markov models in speech recognition,” Foundations and Trends in Signal Processing , vol. 1, no. 3, pp. 195–304, 2008
2008
Earlier work this paper cites.
D. Povey, D. Kanevsky, B. Kingsbury, B. Ramabhadran, G. Saon, and K. Visweswariah, “Boosted mmi for model and feature-space discriminative training,” in ICASSP . IEEE, 2008, pp. 4057–4060
2008
Earlier work this paper cites.
R. Raina, A. Madhavan, and A. Y. Ng, “Large-scale deep unsupervised learning using graphics processors.” in ICML , vol. 9, 2009, pp. 873–880
2009
Cited alongside, same era.
H. Lee, R. Grosse, R. Ranganath, and A. Y. Ng, “Convolutional deep belief networks for scalable unsupervised learning of hierarchical representations,” in ICML . ACM, 2009, pp. 609–616
2009
Cited alongside, same era.
P. Vincent, H. Larochelle, I. Lajoie, Y. Bengio, and P. A. Manzagol, “Stacked denoising autoencoders: Learning useful representations in a deep network with a local denoising criterion,” The Journal of Machine Learning Research , vol. 11, pp. 3371–3408, 2010
2010
Cited alongside, same era.
A. Mohamed, G. Dahl, and G. Hinton, “Acoustic modeling using deep belief networks,” Audio, Speech, and Language Processing, IEEE Transactions on , no. 99, 2010
2010
Cited alongside, same era.
G. Dahl, T. Sainath, and G. Hinton, “Improving Deep Neural Networks for LVCSR using Rectified Linear Units and Dropout,” in ICASSP , 2013
2013
Later among the works it cites.
M. Zeiler, M. Ranzato, R. Monga, M. Mao, K. Yang, Q. Le, P. Nguyen, A. Senior, V. Vanhoucke, J. Dean, and G. Hinton, “On Rectified Linear Units for Speech Processing,” in ICASSP , 2013
2013
Later among the works it cites.
A. Maas, A. Hannun, and A. Ng, “Rectifier Nonlinearities Improve Neural Network Acoustic Models,” in ICML Workshop on Deep Learning for Audio, Speech, and Language Processing , 2013
2013
Later among the works it cites.
T. Sainath, B. Kingsbury, A. Mohamed, G. Dahl, G. Saon, H. Soltau, T. Beran, A. Aravkin, and B. Ramabhadran, “Improvements to Deep Convolutional Neural Networks for LVCSR,” in ASRU , 2013
2013
Later among the works it cites.
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
J. Martens, “Deep learning via hessian-free optimization,” in ICML , 2010, pp. 735–742
2010
Cited alongside, same era.
G. Dahl, D. Yu, L. Deng, and A. Acero, “Context-Dependent Pre-trained Deep Neural Networks for Large Vocabulary Speech Recognition,” IEEE Transactions on Audio, Speech, and Language Processing , 2011
2011
Cited alongside, same era.
B. Gold, N. Morgan, and D. Ellis, Speech and audio signal processing: processing and perception of speech and music . John Wiley & Sons, 2011
2011
Cited alongside, same era.
G. Dahl, D. Yu, and L. Deng, “Large vocabulary continuous speech recognition with context-dependent DBN-HMMs,” in ICASSP , 2011
2011
Cited alongside, same era.
F. Seide, G. Li, and D. Yu, “Conversational speech transcription using context-dependent deep neural networks.” in Interspeech , 2011, pp. 437–440
2011
Cited alongside, same era.
J. Ngiam, A. Coates, A. Lahiri, B. Prochnow, Q. V. Le, and A. Y. Ng, “On optimization methods for deep learning,” in ICML , 2011, pp. 265–272
2011
Cited alongside, same era.
J. Duchi, E. Hazan, and Y. Singer, “Adaptive subgradient methods for online learning and stochastic optimization,” The Journal of Machine Learning Research , vol. 12, pp. 2121–2159, 2011
2011
Cited alongside, same era.
D. Povey, A. Ghoshal, G. Boulianne, L. Burget, O. Glembek, K. Veselý, N. Goel, M. Hannemann, P. Motlicek, Y. Qian, P. Schwarz, J. Silovsky, and G. Stemmer, “The kaldi speech recognition toolkit,” in ASRU , 2011
2011
Cited alongside, same era.
2013
Later among the works it cites.
L. Deng, G. Hinton, and B. Kingsbury, “New types of deep neural network learning for speech recognition and related applications: An overview,” in ICASSP . IEEE, 2013, pp. 8599–8603
2013
Later among the works it cites.
H. Su, G. Li, D. Yu, and F. Seide, “Error back propagation for sequence training of context-dependent deep networks for conversational speech transcription,” in ICASSP , 2013, pp. 6664–6668
2013
Later among the works it cites.
I. Sutskever, J. Martens, G. Dahl, and G. Hinton, “On the Importance of Momentum and Initialization in Deep Learning,” in ICML , 2013
2013
Later among the works it cites.
A. Coates, B. Huval, T. Wang, D. Wu, A. Ng, and B. Catanzaro, “Deep Learning with COTS HPC Systems,” in ICML , 2013
2013
Later among the works it cites.
I. Sutskever, “Training recurrent neural networks,” Ph.D. dissertation, University of Toronto, 2013
2013
Later among the works it cites.
H. Liao, E. McDermott, and A. Senior, “Large scale deep neural network acoustic modeling with semi-supervised training data for YouTube video transcription,” in ASRU , 2013
2013
Later among the works it cites.
T. Sainath, B. Kingsbury, V. Sindhwani, E. Arisoy, and B. Ramabhadran, “Low-Rank Matrix Factorization for Deep Neural Network Training with High-Dimensional Output Targets,” in ICASSP , 2013
2013
Later among the works it cites.
S. Wager, S. Wang, and P. Liang, “Dropout Training as Adaptive Regularization,” in NIPS , 2013
2013
Later among the works it cites.
A. Saxe, J. McClelland, and S. Ganguli, “Learning Hierarchical Category Structure in Deep Networks,” in CogSci , 2013
2013
Later among the works it cites.
A. Senior, G. Heigold, M. Bacchiani, and H. Liao, “Gmm-free dnn acoustic model training,” in ICASSP . IEEE, 2014, pp. 5602–5606
2014
Closest in time.
T. N. Sainath, B. Kingsbury, G. Saon, H. Soltau, A. rahman Mohamed, G. Dahl, and B. Ramabhadran, “Deep Convolutional Neural Networks for Large-Scale Speech Tasks,” Neural Networks , 2014. [Online]. Available: http://www.sciencedirect.com/science/article/pii/S0893608014002007
2014
Closest in time.
H. Sak, O. Vinyals, G. Heigold, A. Senior, E. McDermott, R. Monga, and M. Mao, “Sequence discriminative distributed training of long short-term memory recurrent neural networks,” in Interspeech , 2014
2014
Closest in time.
H. Sak, A. Senior, and F. Beaufays, “Long short-term memory recurrent neural network architectures for large scale acoustic modeling,” in Interspeech , 2014
2014
Closest in time.
C. Weng, D. Yu, S. Watanabe, and B. Juang, “Recurrent deep neural networks for robust speech recognition,” ICASSP , 2014
2014
Closest in time.
T. Chilimbi, Y. Suzue, J. Apacible, and K. Kalyanaraman, “Project adam: Building an efficient and scalable deep learning training system,” in 11th USENIX Symposium on Operating Systems Design and Implementation , 2014, pp. 571–582
2014
Closest in time.
I. Chung, T. N. Sainath, B. Ramabhadran, M. Picheny, J. Gunnels, V. Austel, U. Chauhari, and B. Kingsbury, “Parallel deep neural network training for big data on blue gene/q,” in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis . IEEE, 2014, pp. 745–753
2014
Closest in time.
2014
Closest in time.
D. Yu and L. Deng, “Deep neural network-hidden markov model hybrid systems,” in Automatic Speech Recognition . Springer, 2015, pp. 99–116
2015
Closest in time.