Fetching the paper…
Reading the bibliography…
Recent studies have shown that deep neural networks (DNNs) perform significantly better than shallow networks and Gaussian mixture models (GMMs) on large vocabulary speech recognition tasks.
“Mean field theory for sigmoid belief networks,”
L. Saul, T. Jaakkola, and M. I. Jordan, · 1996
Earlier work this paper cites.
“Vocal tract length normalization for lvcsr,”
P. Zhan et al., · 1997
Earlier work this paper cites.
Switchboard-1 Release 2
J. Godfrey and E. Holliman, · 1997
Earlier work this paper cites.
“Maximum likelihood linear transformations for HMM-based speech recognition,”
M. J. F. Gales, · 1998
Earlier work this paper cites.
“HMM Adaptation Using Vector Taylor Series for Noisy Speech Recognition,”
A. Acero, L. Deng, T. Kristjansson, and J. Zhang, · 2000
Earlier work this paper cites.
“Minimum phone error and i-smoothing for improved discriminative training,”
D. Povey and P. C. Woodland, · 2002
Earlier work this paper cites.
“Adaptive training with joint uncertainty decoding for robust recognition of noisy data,”
H. Liao and M. J. F. Gales, · 2007
Earlier work this paper cites.
“Using continuous features in the maximum entropy model,”
D. Yu, L. Deng, and A. Acero, · 2009
Cited alongside, same era.
“Discriminative adaptive training with VTS and JUD,”
F. Flego and M. J. F. Gales, · 2009
Cited alongside, same era.
“Roles of pretraining and fine-tuning in context-dependent DBN-HMMs for real-world speech recognition,”
D. Yu, L. Deng, and G. Dahl, · 2010
Cited alongside, same era.
“Noise adaptive training for robust automatic speech recognition,”
O. Kalinli, M. L. Seltzer, J. Droppo, and A. Acero, · 2010
Cited alongside, same era.
“Conversational speech transcription using context-dependent deep neural networks,”
F. Seide, G. Li, and D. Yu, · 2011
Cited alongside, same era.
“Feature engineering in context-dependent deep neural networks for conversational speech transcription,”
F. Seide, G.Li, X. Chen, and D. Yu, · 2011
“Large vocabulary continuous speech recognition with context-dependent dbn-hmms,”
G. E. Dahl, D. Yu, L. Deng, and A. Acero, · 2011
Later among the works it cites.
“Derivative kernels for noise robust ASR,”
A. Ragni and M. J. F. Gales, · 2011
Later among the works it cites.
“Context-dependent pretrained deep neural networks for large vocabulary speech recognition,”
G.E. Dahl, D. Yu, L. Deng, and A. Acero, · 2012
Later among the works it cites.
“Exploiting sparseness in deep neural networks for large vocabulary speech recognition,”
D. Yu, F. Seide, G.Li, and L. Deng, · 2012
Later among the works it cites.
“An application of pretrained deep neural networks to large vocabulary conversational speech recognition,”
N. Jaitly, P. Nguyen, A. Senior, and V. Vanhoucke, · 2012
Later among the works it cites.
“Improving wideband speech recognition using mixed-bandwidth training data in CD-DNN-HMM,”
alphaXiv searches the wider corpus for related work and actual follow-ups.
alphaXiv is searching for related work…
Cited alongside, same era.
“Making deep belief networks effective for large vocabulary continuous speech recognition,”
T. N. Sainath, B. Kingsbury, B. Ramabhadran, P. Fousek, P. Novak, and A. r. Mohamed, · 2011
Cited alongside, same era.
“Maximum mutual information estimation of hidden markov model parameters for speech recognition,”
L. Bahl, P. Brown, P.V. De Souza, and R. Mercer,
Cited in the paper.
“Aurora working group: DSR front end LVCSR evaluation AU/384/02,”
N. Parihar and J. Picone,
Cited in the paper.
J. Li, D. Yu, J.-T. Huang, and Y. Gong, · 2012
Later among the works it cites.
“Speaker and noise factorisation for robust speech recognition,”
Y.-Q. Wang and M. J. F. Gales, · 2012
Later among the works it cites.